From Tool to Infrastructure: The Digital Servitization of ATR with eScriptorium
This paper presents the work carried out in the context of H2IOSC, a NextGenerationEU-funded research project. The goal is to provide a complete set of APIs that can facilitate the use of eScriptorium in automated research workflows. All this was done without changing the source code of this open-source platform. eScriptorium is currently designed to be mainly used manually via its web interface. Its RESTful APIs, while present, are only partially documented. The tool we are presenting here is the eScriptorium Proxy API. It has been designed to simplify batch processing of large multipage collections and to enable seamless integration with external portals, code notebooks, and workflow engines by exposing asynchronous REST endpoints that orchestrate the complete Automatic Text Recognition pipeline: image upload, document creation, page import, layout segmentation, text recognition, export, and retrieval of consolidated results. Long-running tasks are executed in the background, allowing clients to submit jobs and monitor their status without blocking their own operations. The service has become a valuable asset of the H2IOSC ecosystem: it has already been integrated with some platforms to allow the automatic transcription of digitized texts and medieval manuscripts. Future integrations may include the Language Resource Switchboard, a service of the CLARIN Research Infrastructure that helps researchers to find and use text processing tools. The Proxy API is currently accessible in its staging environment and is in the process of being deployed in production at the new CNR-ILC server farm: its code has been released under the GPLv3 open-source license.
Authors
- Luca De Santis (ORCID: https://orcid.org/0000-0003-0527-840X)
- Angelo Mario Del Grosso (ORCID: https://orcid.org/0000-0002-4867-6304)
- Federico Boschetti (ORCID: https://orcid.org/0000-0002-7810-7735)
- Chiara Aiola (ORCID: https://orcid.org/0009-0000-4792-4612)
- Monica Monachini (ORCID: https://orcid.org/0000-0003-3356-3988)
- Nicola Baglini (ORCID: https://orcid.org/0009-0001-2650-7745)
Institutions
- Net7 (Italy) (IT)
- Institute for Computational Linguistics “A. Zampolli” (IT)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-24
- DOI
- https://doi.org/10.5281/zenodo.22940497
- Primary Topic
- Digital Humanities and Scholarship
- Type
- article
- Field-Weighted Citation Impact
- 0.00