Trillion-scale integrated framework for high-throughput materials databases and seamless sharing
Abstract Artificial intelligence is accelerating materials discovery, yet for complex systems such as high-entropy alloys, exhaustive exploration of their immense composition spaces is computationally intractable. Here we propose an integrated framework that enables large-scale exploration with reduced resource demands. Our framework uses stateless parallel computing with message queues and containerization, and a lakehouse architecture with columnar storage and metadata-based cloud sharing. Using this framework, we constructed a ten-trillion-scale database of six-principal-element high-entropy alloys in one week, totaling 17.5 terabytes and encompassing compositions, descriptors, and predicted properties. Compared with distributed frameworks, our approach cuts runtime-memory cost by 45%, storage by 90% versus row-based formats, and accelerates subset queries up to 180 times. A full database search completes in under eight minutes using less than four gigabytes of memory. Screening over 50 billion compositions with experimental validation across two alloy families confirms the practical utility of this framework for data-driven materials discovery.
Authors
- Hanwei Fu (ORCID: https://orcid.org/0000-0001-8958-9648)
- Shuo Shi
- Xiaoya Huang
- Yu Liu
- Zengzeng Liang
- Miao Zhou
- Lei Zheng
- Yuanyuan Zhang
- Peng Kang
Institutions
- Tianmushan Laboratory (CN)
- Beihang University (CN)
Publication Details
- Journal
- Nature Communications
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s41467-026-78223-3
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00