Trillion-scale integrated framework for high-throughput materials databases and seamless sharing

Abstract Artificial intelligence is accelerating materials discovery, yet for complex systems such as high-entropy alloys, exhaustive exploration of their immense composition spaces is computationally intractable. Here we propose an integrated framework that enables large-scale exploration with reduced resource demands. Our framework uses stateless parallel computing with message queues and containerization, and a lakehouse architecture with columnar storage and metadata-based cloud sharing. Using this framework, we constructed a ten-trillion-scale database of six-principal-element high-entropy alloys in one week, totaling 17.5 terabytes and encompassing compositions, descriptors, and predicted properties. Compared with distributed frameworks, our approach cuts runtime-memory cost by 45%, storage by 90% versus row-based formats, and accelerates subset queries up to 180 times. A full database search completes in under eight minutes using less than four gigabytes of memory. Screening over 50 billion compositions with experimental validation across two alloy families confirms the practical utility of this framework for data-driven materials discovery.

Authors

Institutions

Publication Details

Journal
Nature Communications
Published
2026-10-07
DOI
https://doi.org/10.1038/s41467-026-78223-3
Primary Topic
Machine Learning in Materials Science
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Trillion-scale integrated framework for high-throughput materials databases and seamless sharing

Hanwei Fu, Shuo Shi, Xiaoya Huang, Yu Liu et al.
Nature Communications
Machine Learning in Materials Science
article

Trillion-scale integrated framework for high-throughput materials databases and seamless sharing

Hanwei Fu, Shuo Shi, Xiaoya Huang, Yu Liu, Zengzeng Liang, Miao Zhou, Lei Zheng, Yuanyuan Zhang, Peng Kang
article en

Abstract

Abstract Artificial intelligence is accelerating materials discovery, yet for complex systems such as high-entropy alloys, exhaustive exploration of their immense composition spaces is computationally intractable. Here we propose an integrated framework that enables large-scale exploration with reduced resource demands. Our framework uses stateless parallel computing with message queues and containerization, and a lakehouse architecture with columnar storage and metadata-based cloud sharing. Using this framework, we constructed a ten-trillion-scale database of six-principal-element high-entropy alloys in one week, totaling 17.5 terabytes and encompassing compositions, descriptors, and predicted properties. Compared with distributed frameworks, our approach cuts runtime-memory cost by 45%, storage by 90% versus row-based formats, and accelerates subset queries up to 180 times. A full database search completes in under eight minutes using less than four gigabytes of memory. Screening over 50 billion compositions with experimental validation across two alloy families confirms the practical utility of this framework for data-driven materials discovery.

Nature Communications
Tianmushan Laboratory (CN), Beihang University (CN)
Openalex Percentile: Top 44%
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.