Fine-Tuning DeepSeek-OCR-2 for Molecular Structure Recognition
Abstract Optical chemical structure recognition (OCSR) is critical for converting 2D molecular diagrams from printed literature into machine-readable formats. While vision-language models have shown promise in end-to-end OCR tasks, their direct application to OCSR remains challenging, and direct full-parameter supervised fine-tuning often fails. In this work, we adapt DeepSeek-OCR-2 for molecular optical recognition by formulating the task as an image-conditioned SMILES generation. To overcome training instabilities, we propose a two-stage progressive supervised fine-tuning strategy: starting with parameter-efficient LoRA and transitioning to selective full-parameter fine-tuning with split learning rates. We train our model on a large-scale corpus combining synthetic renderings from PubChem and realistic patent images from USPTO-MOL to improve the coverage and robustness. Our fine-tuned model, MolSeek-OCR, demonstrates competitive capabilities, achieving exact-match accuracies comparable to those of the best-performing image-to-sequence model. However, it remains inferior to state-of-the-art image-to-graph models. Exploratory post-training shows mixed results: ReFT provides small benchmark-dependent gains, whereas GSPO does not yield stable improvement and can collapse under extended optimization.
Authors
- Haocheng Tang (ORCID: https://orcid.org/0009-0001-5702-2847)
- Junmei Wang (ORCID: https://orcid.org/0000-0002-9607-8229)
- Xingyu Dang
Institutions
- University of Pittsburgh (US)
- Princeton University (US)
Publication Details
- Journal
- ACS Omega
- Published
- 2026-09-11
- DOI
- https://doi.org/10.1021/acsomega.6c04159
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Institute of General Medical Sciences
- Division of Chemistry