Can Small Language Models Learn Materials Synthesis From Limited Experiments? A ZnO‐to‐ZIF‐8 Case Study
We investigate whether small language models can learn synthesis–outcome relationships from a limited experimental dataset using ZnO‐to‐ZIF‐8 conversion as a case study. Several small language models are benchmarked against conventional machine‐learning methods under the same cross‐validation protocol. The best‐performing models, TinyLlama 1.1B and Qwen2, achieve an accuracy of approximately 0.93, comparable to that of the best conventional machine‐learning models for this classification task. We also find that the prediction accuracy remains above 0.8 when each training split contains about 45 samples. This indicates predictive capability under limited‐data conditions. A solvent‐removal analysis further shows that the TinyLlama predictions are only weakly sensitive to solvent‐name tokens. Our study reveals that small language models can perform comparably to conventional machine‐learning methods even in a limited‐data setting and open a new direction for language‐model‐based prediction in experimental materials synthesis.
Authors
- Anh Duc Phan (ORCID: https://orcid.org/0000-0002-8667-1299)
- Ngo T. Que (ORCID: https://orcid.org/0009-0006-6624-6398)
- Nguyen T. T. Duyen
- Nguyen A. Duc
Institutions
- Phenikaa University (VN)
- VinUniversity (VN)
Publication Details
- Journal
- physica status solidi (b)
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1002/pssb.70323
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00