"Signal Certification in the Voynich Manuscript: A Statistical Auditing Framework with Simpson Paradox Resolution and Artifact Probability Scoring

AbstractPurpose. To develop and apply a pre-registered statistical framework distinguishing genuine structuralsignal from analytical artifact in the Voynich Manuscript (Yale Beinecke MS 408), a corpus whosesemantics remain unknown.Design/methodology/approach. A three-tier evidentiary architecture was applied to 27,137 tokens(Langue A: n=6,701; Langue B: n=20,436, Currier classification): (i) encoding-independent structural tests(rarefaction, burstiness, cross-corpus triangulation), (ii) a positional numerical audit using one illustrativeencoding (GC9), tested for sensitivity to the choice of modulus, and (iii) a formal artifact-certificationlayer (Simpson-paradox diagnostics, a 37-system robustness battery, and a five-component ArtifactProbability Score). All hypothesis classes were pre-registered and Bonferroni-corrected.Findings. Two classes of result emerged with different evidentiary status. First, three structuralproperties hold independently of numerical encoding and are the paper's most defensible findings: lexicaldiversity exceeds all nine reference corpora (100th percentile); burstiness analysis corroborates lexicalstructure beyond first-order Markov dynamics; and cross-corpus triangulation (Copiale cipher,Clementine Vulgate) situates this diversity as anomalous even against independent ciphertext andnatural-language controls. Second, two positional signals specific to Langue B (GC9=33 START,p=8.6×10−6; GC9=26 MIDDLE, p=8.3×10−13) are robust within the GC9 encoding but did not replicateunder the alternative moduli individually tested (5, 7, 11); a blind scan across 14 moduli confirms,however, that position-dependent structure itself is genuine and encoding-independent, so only thespecific values 33/26 remain encoding-conditional. A corpus-level anticorrelation, initially suspected toreflect a sign-reversing Simpson's paradox, was audited and is better characterized as a weakcompositional confound (APS=62.6) without a classic sign reversal. Results overall are most compatiblewith natural language or complex cipher, least compatible with hoax construction.Originality. This study introduces the Artifact Probability Score (APS), a five-component composite indexfor certifying statistical signals in corpora of unknown semantics, empirically calibrated against nine external anchors spanning six languages, explicitly separates encoding-independent from encoding-conditional claims, and models transparent null-result and post-hoc-correction reporting throughout. Contribution to the field of Digital Humanities. The framework generalizes beyond the VoynichManuscript to any undeciphered or non-semantic corpus, offering digital humanists a replicablemethodology for signal certification under representation-dependence — and a worked example of what changes, and what doesn't, when a reproducibility audit is taken seriously. Full reproducibility (seed-fixed, containerized, unit-tested) is provided via OSF.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23191817
Primary Topic
Digital Humanities and Scholarship
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

"Signal Certification in the Voynich Manuscript: A Statistical Auditing Framework with Simpson Paradox Resolution and Artifact Probability Scoring

Ahmed Benseddik
Zenodo (CERN European Organization for Nuclear Research)
Digital Humanities and Scholarship
preprint

"Signal Certification in the Voynich Manuscript: A Statistical Auditing Framework with Simpson Paradox Resolution and Artifact Probability Scoring

Ahmed Benseddik
preprint en

Abstract

AbstractPurpose. To develop and apply a pre-registered statistical framework distinguishing genuine structuralsignal from analytical artifact in the Voynich Manuscript (Yale Beinecke MS 408), a corpus whosesemantics remain unknown.Design/methodology/approach. A three-tier evidentiary architecture was applied to 27,137 tokens(Langue A: n=6,701; Langue B: n=20,436, Currier classification): (i) encoding-independent structural tests(rarefaction, burstiness, cross-corpus triangulation), (ii) a positional numerical audit using one illustrativeencoding (GC9), tested for sensitivity to the choice of modulus, and (iii) a formal artifact-certificationlayer (Simpson-paradox diagnostics, a 37-system robustness battery, and a five-component ArtifactProbability Score). All hypothesis classes were pre-registered and Bonferroni-corrected.Findings. Two classes of result emerged with different evidentiary status. First, three structuralproperties hold independently of numerical encoding and are the paper's most defensible findings: lexicaldiversity exceeds all nine reference corpora (100th percentile); burstiness analysis corroborates lexicalstructure beyond first-order Markov dynamics; and cross-corpus triangulation (Copiale cipher,Clementine Vulgate) situates this diversity as anomalous even against independent ciphertext andnatural-language controls. Second, two positional signals specific to Langue B (GC9=33 START,p=8.6×10−6; GC9=26 MIDDLE, p=8.3×10−13) are robust within the GC9 encoding but did not replicateunder the alternative moduli individually tested (5, 7, 11); a blind scan across 14 moduli confirms,however, that position-dependent structure itself is genuine and encoding-independent, so only thespecific values 33/26 remain encoding-conditional. A corpus-level anticorrelation, initially suspected toreflect a sign-reversing Simpson's paradox, was audited and is better characterized as a weakcompositional confound (APS=62.6) without a classic sign reversal. Results overall are most compatiblewith natural language or complex cipher, least compatible with hoax construction.Originality. This study introduces the Artifact Probability Score (APS), a five-component composite indexfor certifying statistical signals in corpora of unknown semantics, empirically calibrated against nine external anchors spanning six languages, explicitly separates encoding-independent from encoding-conditional claims, and models transparent null-result and post-hoc-correction reporting throughout. Contribution to the field of Digital Humanities. The framework generalizes beyond the VoynichManuscript to any undeciphered or non-semantic corpus, offering digital humanists a replicablemethodology for signal certification under representation-dependence — and a worked example of what changes, and what doesn't, when a reproducibility audit is taken seriously. Full reproducibility (seed-fixed, containerized, unit-tested) is provided via OSF.

Zenodo (CERN European Organization for Nuclear Research)
Digital Humanities and Scholarship
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.