ChemE-MTDS:Multi-Turn Dialogue Data Synthesis and Quality-Controlled Fine-Tuning for Chemical Engineering Large Language Models
Large language models (LLMs) have demonstrated remarkable capabilities in natural language understanding and reasoning. However, their performance in multi-turn dialogues within specialized domains, such as chemical engineering, remains limited due to the lack of high-quality, context-rich instruction datasets. In this paper, we propose ChemE-MTDS, a multi-turn dialogue data synthesis and quality-controlled fine-tuning framework specifically tailored for the chemical engineering domain. The framework integrates four complementary Question–Answering (QA) synthesis strategies: Continuous Scenario QA, Error Correction QA, Style Diverse QA, and Context Independent QA. To ensure data reliability, we design a multi-dimensional automated quality control pipeline that evaluates contextual memory, anaphora resolution, and logical coherence. Fine-tuning the Spark-70B model with the synthesized dataset demonstrates substantial improvements over baseline models in contextual understanding, cross-turn information integration, and logical consistency. Moreover, ChemE-Spark-70B surpasses general-purpose LLMs, highlighting the importance of domain-specific multi-turn dialogue data. This study provides a practical methodology for enhancing reasoning and multi-turn interaction capabilities in specialized domains.
Authors
- Xin Li (ORCID: https://orcid.org/0000-0001-8696-5361)
- Feiyang Xu
- Le Wu
- Defu Lian
- Yi Li
- Mingxin Miao
Institutions
- University of Science and Technology of China (CN)
- Hefei University of Technology (CN)
- IFlyTek (China)
Publication Details
- Journal
- International Journal of Computational Intelligence Systems
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1007/s44196-026-01597-1
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00