Unravelling dataset effects in machine-learning screening of perovskite oxides for solar hydrogen production
Machine-learning (ML) models are increasingly used to accelerate the discovery of oxide materials for solar thermochemical hydrogen production (STCH), where oxygen vacancy formation enthalpy ( ∆ h o ) is a key descriptor of redox behaviour. However, the predictive reliability of ∆ h o depends on the choice of regression model and on how the underlying dataset defines the learning problem, including its compositional diversity, oxygen non-stoichiometry, and thermodynamic definition of the target. This aspect has commonly been overlooked in previous studies. We develop an ML workflow that explicitly evaluates dataset effects in the prediction of Δ h o . Rather than treating available data as a single source, three independent density functional theory (DFT) datasets were analysed as distinct learning problems, as they differ in compositional space, oxygen non-stoichiometry, and computational treatment of reduction, including phase transitions. This work shows how feature selection (FS) is strongly dependent on the dataset, although recurrent B-site-related descriptors emerge across all cases, supporting a consistent physicochemical interpretation. The highest predictive accuracy is obtained for the largest and most stoichiometrically homogeneous dataset, whereas the lowest extrapolation risk is found in the dataset that more evenly spans the Δ h o , the compositional, and stoichiometric space. External validation further showed that the model trained on the dataset with greater physical and chemical representativeness achieved the closest agreement with reference values. Thus, this work shows that high accuracy and low mean absolute error do not necessarily imply reliable prediction across a broad discovery space; other factors such as extrapolation risk and external validation are essential.
Authors
- Estefanía Fernández-Villanueva (ORCID: https://orcid.org/0000-0002-9419-0786)
- Silvia Jiménez-Fernández (ORCID: https://orcid.org/0000-0002-2065-1754)
- Laura Molina (ORCID: https://orcid.org/0009-0006-2367-611X)
- Alicia Bayón (ORCID: https://orcid.org/0000-0002-6422-0884)
- M. Verónica Ganduglia-Pirovano
- Khalid Achbab-Gueriguer
- Alberto de la Calle
Institutions
- Instituto de Catálisis y Petroleoquímica (ES)
- Universidad Autónoma de Madrid (ES)
- Universidad Politécnica de Madrid (ES)
Publication Details
- Journal
- Computational Materials Science
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1016/j.commatsci.2026.115111
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00