Turning Scanned Classical Texts into Searchable Markdown
A practical study of low-cost vision models, how they fail, and a self-correcting pipeline that makes them trustworthy Working paper-September 2026 Altaf Ali, independent researcher-with an AI coding agent (DeepSeek Harness running DeepSeek-V4-Flash), which wrote the application, designed and ran the experiments, and analysed the results across iterative sessions. Transcription models compared: DeepSeek V4-Flash Vision (deepseek-flash) and Google Gemini 3.1 Flash Lite. All measurements come from real books processed on a local Mac workstation. 1. Why this tool is needed 1.1 The people and the texts An independent researcher, a seminary student, or a graduate student in Islamic studies typically works with printed editions of classical Arabic, Persian (Farsi) and Urdu works: Qurʾānic commentary (tafsīr), ḥadīth collections, theology, philosophy, law, poetry. The copies they own are almost always scanned page images-photographs of paper. There is usually no text layer underneath, and the text cannot be selected, copied, or searched.
Authors
- Altaf Ali
Publication Details
- Journal
- Knowledge Commons (Lakehead University)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.17613/gmcwa-s7j89
- Primary Topic
- Digital Humanities and Scholarship
- Type
- article
- Field-Weighted Citation Impact
- 0.00