Structural Decipherment Without Lexical Recovery: A Cross-Validated Morphological Analysis of the Voynich Manuscript
The Voynich Manuscript exhibits strong statistical regularities, but the distinction between recoverable grammar and recoverable lexical meaning remains unresolved. This study asks how far a progressively frozen, cross-validated analysis can reconstruct the manuscript’s internal word-building system without assigning meanings from illustrations, candidate languages, or post-hoc English paraphrase.Using canonical transcription, protected holdout partitions, alternate-transcription replication, exact host substitutions, and adversarial stopping rules, the analysis recovers a cumulative left-edge morphology, a transferable inventory of 39 surface CORE-0 stem/bodies, reproducible host-by-right-state coupling, a stable K/T referential contrast, and a directed network among powered lexical cores.The 39-core inventory predicts protected-frame combinatorics and content-blind contextual fingerprints out of sample, while attempts to decompose these bodies into a deeper transferable root layer fail. Nine recurring hosts show strongly replicated coupling to EY/EEY/EDY right-edge states, supporting a distributed morphological paradigm rather than independent left- and right-side systems. A 28-node directed CORE network also replicates across canonical calibration, untouched Holdout, and alternate-transcription tests, yielding neutral outbound- and inbound-biased stem classes without requiring part-of-speech or semantic assignments.On the project’s fully constrained 1,609-line benchmark, 12,299 of 12,495 positions are role-naturalizable (98.4314%), leaving 196 opaque positions. This percentage is explicitly a conditional benchmark-maturity statistic, not a measure of literal plaintext recovery or manuscript-wide translation accuracy.Repeated attempts to cross from structural organization to ordinary lexical meaning—including deeper-root induction, topic clustering, functional-neighborhood anchoring, local lexical syntax, illustration matching, source-language fitting, and prospective semantic transfer—fail the project’s promotion criteria. The resulting contribution is therefore a structural decipherment rather than a lexical one: a reproducible grammatical and morphological architecture together with an empirically identified lexical-identifiability ceiling.The work extends, but is distinct from, the earlier controlled-translation release at DOI 10.5281/zenodo.22049207.
Authors
- Matthew Dominik (ORCID: https://orcid.org/0009-0007-9449-2630)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23113674
- Primary Topic
- Intelligence, Security, War Strategy
- Type
- preprint