Leibniz Legible: an open access layer for the digitized Leibniz Nachlass (project statement, September 2026)

Leibniz Legible: an open access layer for the digitized Leibniz Nachlass (project statement, September 2026) — a project statement of the Leibniz Legible project. Abstract. The Leibniz Nachlass in Hannover, some 236,000 digitized page images, is the largest body of writing by any early-modern philosopher that remains mostly unread: after a century of work the Akademie-Ausgabe has printed roughly a quarter of it. Leibniz Legible is a one-person open project that builds an access layer over those images rather than an edition: every page machine-transcribed with a per-line confidence score and a full provenance record, searchable with tolerance for early-modern orthography, and browsable in an open IIIF viewer that loads the library's own images and shows the machine reading beside them with honest labels. This statement records the state of the project in September 2026. The digitized corpus has been counted for the first time from the library's own metadata (2,225 works, 236,795 page images). The character error rate of the ERC PHILIUMM project's Leibniz handwriting model has been reproduced under a frozen protocol (7.95%, 95% CI 7.49–8.46, against a claimed 8.33%), and the same protocol gives the first published numbers for frontier vision-language models on Leibniz's hand (37% to 79% character error rate, zero-shot). The whole corpus has been segmented and read with that model: 236,210 pages, 13.5 million lines, each with a confidence and a run record, on one desktop computer in six weeks. A ground-truth factory has aligned the reading text of copyright-expired edition volumes back onto the machine lines and minted 297,424 training pairs, six times the target, at a precision that is measured on favourable material but only preliminarily audited on the corpus. The serving layer (a search index, a JSON API, IIIF Presentation 3 manifests with W3C annotations, and a viewer) is built, and the datasets are packaged with cards that carry their provenance, licence and error rates. The statement closes with what the next model can and cannot buy, what the crowd and language models can and cannot fix, and the roadmap to the public release. Provenance. Generated from the project repository https://github.com/marchofhares/leibnizlegible at commit ede16e1; licensed CC BY 4.0. Images are the GWLB's (Public Domain Mark 1.0, never rehosted); the catalogue is the Leibniz-Katalog (BBAW/TELOTA, CC BY 4.0); the HTR model and ground truth are from the ERC project PHILIUMM (CC BY 4.0).

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22782813
Primary Topic
Digital Humanities and Scholarship
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Leibniz Legible: an open access layer for the digitized Leibniz Nachlass (project statement, September 2026)

Evan Tabak Atlas
Zenodo (CERN European Organization for Nuclear Research)
Digital Humanities and Scholarship
article

Leibniz Legible: an open access layer for the digitized Leibniz Nachlass (project statement, September 2026)

Evan Tabak Atlas
article en

Abstract

Leibniz Legible: an open access layer for the digitized Leibniz Nachlass (project statement, September 2026) — a project statement of the Leibniz Legible project. Abstract. The Leibniz Nachlass in Hannover, some 236,000 digitized page images, is the largest body of writing by any early-modern philosopher that remains mostly unread: after a century of work the Akademie-Ausgabe has printed roughly a quarter of it. Leibniz Legible is a one-person open project that builds an access layer over those images rather than an edition: every page machine-transcribed with a per-line confidence score and a full provenance record, searchable with tolerance for early-modern orthography, and browsable in an open IIIF viewer that loads the library's own images and shows the machine reading beside them with honest labels. This statement records the state of the project in September 2026. The digitized corpus has been counted for the first time from the library's own metadata (2,225 works, 236,795 page images). The character error rate of the ERC PHILIUMM project's Leibniz handwriting model has been reproduced under a frozen protocol (7.95%, 95% CI 7.49–8.46, against a claimed 8.33%), and the same protocol gives the first published numbers for frontier vision-language models on Leibniz's hand (37% to 79% character error rate, zero-shot). The whole corpus has been segmented and read with that model: 236,210 pages, 13.5 million lines, each with a confidence and a run record, on one desktop computer in six weeks. A ground-truth factory has aligned the reading text of copyright-expired edition volumes back onto the machine lines and minted 297,424 training pairs, six times the target, at a precision that is measured on favourable material but only preliminarily audited on the corpus. The serving layer (a search index, a JSON API, IIIF Presentation 3 manifests with W3C annotations, and a viewer) is built, and the datasets are packaged with cards that carry their provenance, licence and error rates. The statement closes with what the next model can and cannot buy, what the crowd and language models can and cannot fix, and the roadmap to the public release. Provenance. Generated from the project repository https://github.com/marchofhares/leibnizlegible at commit ede16e1; licensed CC BY 4.0. Images are the GWLB's (Public Domain Mark 1.0, never rehosted); the catalogue is the Leibniz-Katalog (BBAW/TELOTA, CC BY 4.0); the HTR model and ground truth are from the ERC project PHILIUMM (CC BY 4.0).

Zenodo (CERN European Organization for Nuclear Research)
Quality Education
Openalex Percentile: Top 1%
Digital Humanities and Scholarship
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.