Structural Decipherment Without Lexical Recovery: A Cross-Validated Morphological Analysis of the Voynich Manuscript

The Voynich Manuscript exhibits strong statistical regularities, but the distinction between recoverable grammar and recoverable lexical meaning remains unresolved. This study asks how far a progressively frozen, cross-validated analysis can reconstruct the manuscript’s internal word-building system without assigning meanings from illustrations, candidate languages, or post-hoc English paraphrase.Using canonical transcription, protected holdout partitions, alternate-transcription replication, exact host substitutions, and adversarial stopping rules, the analysis recovers a cumulative left-edge morphology, a transferable inventory of 39 surface CORE-0 stem/bodies, reproducible host-by-right-state coupling, a stable K/T referential contrast, and a directed network among powered lexical cores.The 39-core inventory predicts protected-frame combinatorics and content-blind contextual fingerprints out of sample, while attempts to decompose these bodies into a deeper transferable root layer fail. Nine recurring hosts show strongly replicated coupling to EY/EEY/EDY right-edge states, supporting a distributed morphological paradigm rather than independent left- and right-side systems. A 28-node directed CORE network also replicates across canonical calibration, untouched Holdout, and alternate-transcription tests, yielding neutral outbound- and inbound-biased stem classes without requiring part-of-speech or semantic assignments.On the project’s fully constrained 1,609-line benchmark, 12,299 of 12,495 positions are role-naturalizable (98.4314%), leaving 196 opaque positions. This percentage is explicitly a conditional benchmark-maturity statistic, not a measure of literal plaintext recovery or manuscript-wide translation accuracy.Repeated attempts to cross from structural organization to ordinary lexical meaning—including deeper-root induction, topic clustering, functional-neighborhood anchoring, local lexical syntax, illustration matching, source-language fitting, and prospective semantic transfer—fail the project’s promotion criteria. The resulting contribution is therefore a structural decipherment rather than a lexical one: a reproducible grammatical and morphological architecture together with an empirically identified lexical-identifiability ceiling.The work extends, but is distinct from, the earlier controlled-translation release at DOI 10.5281/zenodo.22049207.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23113674
Primary Topic
Intelligence, Security, War Strategy
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Structural Decipherment Without Lexical Recovery: A Cross-Validated Morphological Analysis of the Voynich Manuscript

Matthew Dominik
Zenodo (CERN European Organization for Nuclear Research)
Intelligence, Security, War Strategy
preprint

Structural Decipherment Without Lexical Recovery: A Cross-Validated Morphological Analysis of the Voynich Manuscript

Matthew Dominik
preprint en

Abstract

The Voynich Manuscript exhibits strong statistical regularities, but the distinction between recoverable grammar and recoverable lexical meaning remains unresolved. This study asks how far a progressively frozen, cross-validated analysis can reconstruct the manuscript’s internal word-building system without assigning meanings from illustrations, candidate languages, or post-hoc English paraphrase.Using canonical transcription, protected holdout partitions, alternate-transcription replication, exact host substitutions, and adversarial stopping rules, the analysis recovers a cumulative left-edge morphology, a transferable inventory of 39 surface CORE-0 stem/bodies, reproducible host-by-right-state coupling, a stable K/T referential contrast, and a directed network among powered lexical cores.The 39-core inventory predicts protected-frame combinatorics and content-blind contextual fingerprints out of sample, while attempts to decompose these bodies into a deeper transferable root layer fail. Nine recurring hosts show strongly replicated coupling to EY/EEY/EDY right-edge states, supporting a distributed morphological paradigm rather than independent left- and right-side systems. A 28-node directed CORE network also replicates across canonical calibration, untouched Holdout, and alternate-transcription tests, yielding neutral outbound- and inbound-biased stem classes without requiring part-of-speech or semantic assignments.On the project’s fully constrained 1,609-line benchmark, 12,299 of 12,495 positions are role-naturalizable (98.4314%), leaving 196 opaque positions. This percentage is explicitly a conditional benchmark-maturity statistic, not a measure of literal plaintext recovery or manuscript-wide translation accuracy.Repeated attempts to cross from structural organization to ordinary lexical meaning—including deeper-root induction, topic clustering, functional-neighborhood anchoring, local lexical syntax, illustration matching, source-language fitting, and prospective semantic transfer—fail the project’s promotion criteria. The resulting contribution is therefore a structural decipherment rather than a lexical one: a reproducible grammatical and morphological architecture together with an empirically identified lexical-identifiability ceiling.The work extends, but is distinct from, the earlier controlled-translation release at DOI 10.5281/zenodo.22049207.

Zenodo (CERN European Organization for Nuclear Research)
Intelligence, Security, War Strategy
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.