Fine-Tuning DeepSeek-OCR-2 for Molecular Structure Recognition

Abstract Optical chemical structure recognition (OCSR) is critical for converting 2D molecular diagrams from printed literature into machine-readable formats. While vision-language models have shown promise in end-to-end OCR tasks, their direct application to OCSR remains challenging, and direct full-parameter supervised fine-tuning often fails. In this work, we adapt DeepSeek-OCR-2 for molecular optical recognition by formulating the task as an image-conditioned SMILES generation. To overcome training instabilities, we propose a two-stage progressive supervised fine-tuning strategy: starting with parameter-efficient LoRA and transitioning to selective full-parameter fine-tuning with split learning rates. We train our model on a large-scale corpus combining synthetic renderings from PubChem and realistic patent images from USPTO-MOL to improve the coverage and robustness. Our fine-tuned model, MolSeek-OCR, demonstrates competitive capabilities, achieving exact-match accuracies comparable to those of the best-performing image-to-sequence model. However, it remains inferior to state-of-the-art image-to-graph models. Exploratory post-training shows mixed results: ReFT provides small benchmark-dependent gains, whereas GSPO does not yield stable improvement and can collapse under extended optimization.

Authors

Institutions

Publication Details

Journal
ACS Omega
Published
2026-09-11
DOI
https://doi.org/10.1021/acsomega.6c04159
Primary Topic
Machine Learning in Materials Science
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Fine-Tuning DeepSeek-OCR-2 for Molecular Structure Recognition

Haocheng Tang, Junmei Wang, Xingyu Dang
ACS Omega
Machine Learning in Materials Science
article

Fine-Tuning DeepSeek-OCR-2 for Molecular Structure Recognition

Haocheng Tang, Junmei Wang, Xingyu Dang
article en

Abstract

Abstract Optical chemical structure recognition (OCSR) is critical for converting 2D molecular diagrams from printed literature into machine-readable formats. While vision-language models have shown promise in end-to-end OCR tasks, their direct application to OCSR remains challenging, and direct full-parameter supervised fine-tuning often fails. In this work, we adapt DeepSeek-OCR-2 for molecular optical recognition by formulating the task as an image-conditioned SMILES generation. To overcome training instabilities, we propose a two-stage progressive supervised fine-tuning strategy: starting with parameter-efficient LoRA and transitioning to selective full-parameter fine-tuning with split learning rates. We train our model on a large-scale corpus combining synthetic renderings from PubChem and realistic patent images from USPTO-MOL to improve the coverage and robustness. Our fine-tuned model, MolSeek-OCR, demonstrates competitive capabilities, achieving exact-match accuracies comparable to those of the best-performing image-to-sequence model. However, it remains inferior to state-of-the-art image-to-graph models. Exploratory post-training shows mixed results: ReFT provides small benchmark-dependent gains, whereas GSPO does not yield stable improvement and can collapse under extended optimization.

ACS Omega
University of Pittsburgh (US), Princeton University (US)
National Institute of General Medical Sciences, Division of Chemistry
Quality Education
Openalex Percentile: Top 24%
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Fine-Tuning DeepSeek-OCR-2 for Molecular Structure Recognition — Haocheng Tang, Junmei Wang, et al. · ACS Omega (2026) | TGRS Research Map | TGRS