A Generalizable Framework for Building Executable Domain‐Specific Large Language Models Under Data Scarcity: Demonstration on Semiconductor Technology Computer‐Aided Design Simulation

Scientific and engineering verticals often suffer from data scarcity and strict executability requirements: models must generate not only fluent text but also syntactically valid, tool‐compilable scripts. We present a schema‐first alignment framework for building compact, executable domain‐specific LLMs in low‐resource settings. The framework integrates three core components: (i) large‐scale synthetic QA data generation from expert documentation to instill foundational domain knowledge; (ii) a code‐centric IR DPO workflow that converts verified tool decks into interpretable intermediate representations (IR), performs equivalence‐preserving diversification and constructs preference pairs to directly optimize instruction compliance and code executability; and (iii) a controlled evaluation of retrieval‐augmented generation (RAG), showing that while RAG benefits general LLMs, it can marginally degrade the performance of already domain‐aligned models. We demonstrate the framework by instantiating TcadGPT for semiconductor technology computer‐aided design (TCAD). Using 1.5 M synthetic QA pairs and an IR‐driven DPO dataset, TcadGPT attains 85.6% semantic accuracy and an 85.0% syntax pass rate on SDE executability tests, substantially outperforming state‐of‐the‐art general LLMs such as GPT‐4o. To probe portability beyond TCAD, we apply the same recipe to the open‐source FEM solver Elmer , observing consistent improvements in script‐level success rates over general‐purpose baselines. All datasets, benchmarks, and code (including P1, P2, and IR DPO) are released for reproducibility. Together, these results suggest that the proposed framework provides a robust and reproducible path toward executable LLMs in specialized, data‐scarce professional domains.

Authors

Institutions

Publication Details

Journal
Advanced Intelligent Systems
Published
2026-09-16
DOI
https://doi.org/10.1002/aisy.70520
Primary Topic
Machine Learning in Materials Science
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Generalizable Framework for Building Executable Domain‐Specific Large Language Models Under Data Scarcity: Demonstration on Semiconductor Technology Computer‐Aided Design Simulation

Kai Chang, Zhenhua Wu, Di Wang, Shaohua Wu et al.
Advanced Intelligent Systems
Machine Learning in Materials Science
article

A Generalizable Framework for Building Executable Domain‐Specific Large Language Models Under Data Scarcity: Demonstration on Semiconductor Technology Computer‐Aided Design Simulation

Kai Chang, Zhenhua Wu, Di Wang, Shaohua Wu, Yu Liu
article en

Abstract

Scientific and engineering verticals often suffer from data scarcity and strict executability requirements: models must generate not only fluent text but also syntactically valid, tool‐compilable scripts. We present a schema‐first alignment framework for building compact, executable domain‐specific LLMs in low‐resource settings. The framework integrates three core components: (i) large‐scale synthetic QA data generation from expert documentation to instill foundational domain knowledge; (ii) a code‐centric IR DPO workflow that converts verified tool decks into interpretable intermediate representations (IR), performs equivalence‐preserving diversification and constructs preference pairs to directly optimize instruction compliance and code executability; and (iii) a controlled evaluation of retrieval‐augmented generation (RAG), showing that while RAG benefits general LLMs, it can marginally degrade the performance of already domain‐aligned models. We demonstrate the framework by instantiating TcadGPT for semiconductor technology computer‐aided design (TCAD). Using 1.5 M synthetic QA pairs and an IR‐driven DPO dataset, TcadGPT attains 85.6% semantic accuracy and an 85.0% syntax pass rate on SDE executability tests, substantially outperforming state‐of‐the‐art general LLMs such as GPT‐4o. To probe portability beyond TCAD, we apply the same recipe to the open‐source FEM solver Elmer , observing consistent improvements in script‐level success rates over general‐purpose baselines. All datasets, benchmarks, and code (including P1, P2, and IR DPO) are released for reproducibility. Together, these results suggest that the proposed framework provides a robust and reproducible path toward executable LLMs in specialized, data‐scarce professional domains.

Advanced Intelligent Systems
China Electronic Information Industry Development (CN), Zhejiang Lab (CN), Chemical Industry Press (CN)
Quality Education
Openalex Percentile: Top 24%
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.