Large language model-assisted discovery of dopants for enhanced thermoelectric performance in CoSb₃-based skutterudites
We present a data-driven approach for accelerating the discovery of high-performance CoSb₃-based skutterudites by curating a comprehensive dataset of compositions with various filler elements from over 300 research articles. Leveraging large language models (LLMs), we extract and embed compositional representations, which are then used to train a regression head for predicting thermoelectric figure of merit. Compared to conventional neural-network models using one-hot encoding and composition-based feature vector (CBFV) representations, our LLM-based model achieves lower prediction errors, while demonstrating predictive performance comparable to a CBFV-based random forest model. We further employ the trained model to propose novel filler compositions with promising thermoelectric properties. Finally, we support these predicted candidates through density functional theory and molecular dynamics calculations to assess their electrical and thermal conductivity. This data-driven approach demonstrates the potential of combining natural language processing, machine learning, and quantum simulations for thermoelectric materials design.
Authors
- Houlong L. Zhuang (ORCID: https://orcid.org/0000-0001-7276-7938)
- Yagnik Bandyopadhyay (ORCID: https://orcid.org/0009-0001-9276-1880)
- Dylan Noel Serrao
Institutions
- Arizona State University (US)
Publication Details
- Journal
- Computational Materials Science
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1016/j.commatsci.2026.115143
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00