Impact of descriptor engineering and model selection on regression-based prediction of impurity formation energies in 2D materials

The goal for this task will be to systematically assess the impact of ML, or machine learning workflow components, on the predictive performance of defect energetics for two-dimensional (2D) materials, with a focus on the combined effects of descriptor engineering and model selection on training stability, generalization behaviour and prediction accuracy for impurity formation energy prediction. To achieve this goal, a regression-based ML framework is developed to test statistical learning models and artificial neural network (ANN) architectures. We use a curated database of interstitial and adsorbate impurity configurations in 2D materials to produce two sub-datasets for comparison in distinct defect settings. The engineering of descriptors is done usinvectorized matrix representations of atomic and structural characteristics, and the performance of the models are assessed using tree-based ensemble regressors and ANN models. We find that a descriptor transformation significantly improves the predictive performance of all models, with statistical learning methods achieving training errors below 1.4 eV (interstitial dataset) and 1.1 eV (adsorbate dataset), while the errors obtained with the Artificial Neural Networks (ANN) models are 2.1 eV and 1.3 eV, respectively. Testing on unobserved data shows better generalization with engineered descriptors, leading to lower prediction errors across both model classes. However, residual overfitting is still more pronounced on the interstitial dataset than on the adsorbate systems, indicating that ML performance is sensitive to the complexity and heterogeneity of the defect dataset. This study systematically benchmarks ML processes for materials property prediction and offers guidance on data engineering and model selection for regression-based materials informatics.

Authors

Institutions

Publication Details

Journal
Computational Materials Science
Published
2026-09-30
DOI
https://doi.org/10.1016/j.commatsci.2026.114993
Primary Topic
Machine Learning in Materials Science
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Impact of descriptor engineering and model selection on regression-based prediction of impurity formation energies in 2D materials

Lin Kong
Computational Materials Science
Machine Learning in Materials Science
article

Impact of descriptor engineering and model selection on regression-based prediction of impurity formation energies in 2D materials

Lin Kong
article en

Abstract

The goal for this task will be to systematically assess the impact of ML, or machine learning workflow components, on the predictive performance of defect energetics for two-dimensional (2D) materials, with a focus on the combined effects of descriptor engineering and model selection on training stability, generalization behaviour and prediction accuracy for impurity formation energy prediction. To achieve this goal, a regression-based ML framework is developed to test statistical learning models and artificial neural network (ANN) architectures. We use a curated database of interstitial and adsorbate impurity configurations in 2D materials to produce two sub-datasets for comparison in distinct defect settings. The engineering of descriptors is done usinvectorized matrix representations of atomic and structural characteristics, and the performance of the models are assessed using tree-based ensemble regressors and ANN models. We find that a descriptor transformation significantly improves the predictive performance of all models, with statistical learning methods achieving training errors below 1.4 eV (interstitial dataset) and 1.1 eV (adsorbate dataset), while the errors obtained with the Artificial Neural Networks (ANN) models are 2.1 eV and 1.3 eV, respectively. Testing on unobserved data shows better generalization with engineered descriptors, leading to lower prediction errors across both model classes. However, residual overfitting is still more pronounced on the interstitial dataset than on the adsorbate systems, indicating that ML performance is sensitive to the complexity and heterogeneity of the defect dataset. This study systematically benchmarks ML processes for materials property prediction and offers guidance on data engineering and model selection for regression-based materials informatics.

Computational Materials ScienceVol. 276
Xiamen University (CN)
Affordable and clean energy
Openalex Percentile: Top 26%
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Impact of descriptor engineering and model selection on regression-based prediction of impurity formation energies in 2D materials — Lin Kong · Computational Materials Science (2026) | TGRS Research Map | TGRS