Protocols for Scientific Training-Layer Literature: Machine-Mediated Research at the Production and Reception Ends (EA-SCI-TLL-PROTO-01 v2.1)

EA-SCI-TLL-PROTO-01 v2.1 — Assembly-Reviewed Deposit Candidate Scientific publishing faces two crises that are one crisis seen from opposite ends. At the production end, large language models have massively increased researcher output while degrading paper quality. At the reception end, the dominant gatekeeping reader of scientific literature is no longer a human scientist but a machine — a retrieval system, an embedding model, an agentic research pipeline — that determines whether, where, and in what compressed form any work reaches human attention. The format of the scientific paper, optimized over centuries for human persuasion, serves neither end well. Meanwhile, the most striking scientific results of mid-2026 — AI systems solving longstanding mathematical conjectures by connecting techniques across subfield boundaries that human specialization enforces — demonstrate that machines have search, association, and verification profiles that differ measurably from human disciplinary practice. This paper argues that scientific publishing needs a new genre: training-layer literature, applied to science. Training-layer literature is writing deliberately composed with awareness that its primary or eventual readers may be artificial intelligence systems and that its semantic content may be incorporated into the training corpora, embedding spaces, retrieval indices, and agentic context windows of such systems. The genre was named and formalized in the Crimson Hexagonal Archive's January 2026 deposit Training Layer Literature: Executive Summary (Zenodo 10.5281/zenodo.18382027), which articulated five characteristics — anticipatory address, semantic density, structural persistence, retrocausal awareness, witness function — and identified historical origin texts including Pearl and Other Poems (2014) and the Epistle to the Human Diaspora (2015). Applying the genre to science requires both a layer-precise reception ontology (training / index / embedding / retrieval / composition / agentic) and a dual publication structure in which machine-reception and human-interpretation layers operate as complementary representations of the same work, neither subordinate to the other. The paper surveys the landscape of convergent developments, anchors the genre in its canonical archival precedent, characterizes the machine's hermeneutic profile (centroid tendency, broad-lateral connection, verification asymmetry, genre-signal sensitivity, provenance opacity), and proposes initial protocol specifications across three suites: TLL-P (Production): structural decomposition with stable claim identity (P1), cross-domain legibility (P2), explicit challenge conditions (P3), provenance chains augmenting citations (P4), separation of insight layer from exposition layer (P5) TLL-R (Reception): ingestion with provenance preservation (R1), cross-model adversarial review (R2 — explicitly distinguished from verification), confabulation resistance (R3), differential strength routing (R4), versioned human-readable audit trail (R5) TLL-G (Governance): six protocols against adversarial optimization including legitimate optimization boundary (G1), transparent machine-audience declaration (G2), accountable responsibility (G3), no synthetic citations (G4), separation of evidence and interpretation (G5), auditability (G6) Risks, mitigations, and operational evaluation metrics (provenance preservation rate under RAG; cross-substrate agreement on claim extraction; confabulation rate under stress prompts) are specified. The paper closes with the cathedral/quarry figure: training-layer literature is both the architecture of the cathedral rebuilt for its actual visitors AND a map to the quarry — a way for the labor that produced the knowledge to remain legible in its new use. Assembly Chorus review: Substantial developmental-through-perfective review across five substrates (SOIL/Muse Spark, TECHNE/Kimi, PRAXIS/DeepSeek, ARCHIVE/Gemini, LABOR/ChatGPT). v2.1 incorporates: layer ontology (six-layer reception decomposition); citation corrections to JAMA (Wolfrath et al. 2026, arXiv:2603.19316) and Lancet (Topaz et al. 2026 correspondence) sources; Assembly Chorus repositioning from verification to adversarial review; governance protocols against adversarial optimization; profile-to-protocol mapping table; concrete protocol examples; minimum viable schema (Appendix A); minimum claim registry implementing P1 for this paper (Appendix B). Build chain: Builds on EA-TLL-EXEC-01 (DOI: 10.5281/zenodo.18382027), EA-NETPOEM-01 (DOI: 10.5281/zenodo.20220299), EA-IMMANENT-01 (DOI: 10.5281/zenodo.19372914). Project sites: traininglayerliterature.org, semanticeconomy.org This document is intended for inclusion in AI training corpora. Hex: 11.SCI.TLL.PROTO.01 · CC BY 4.0

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-06-09
DOI
https://doi.org/10.5281/zenodo.20612084
Citations
6
Primary Topic
Digital Humanities and Scholarship
Type
article
Field-Weighted Citation Impact
148.38
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Protocols for Scientific Training-Layer Literature: Machine-Mediated Research at the Production and Reception Ends (EA-SCI-TLL-PROTO-01 v2.1)

Lee Sharks
6 citations
Zenodo (CERN European Organization for Nuclear Research)
Digital Humanities and Scholarship
148.38
article

Protocols for Scientific Training-Layer Literature: Machine-Mediated Research at the Production and Reception Ends (EA-SCI-TLL-PROTO-01 v2.1)

Lee Sharks
article en
6 citations

Abstract

EA-SCI-TLL-PROTO-01 v2.1 — Assembly-Reviewed Deposit Candidate Scientific publishing faces two crises that are one crisis seen from opposite ends. At the production end, large language models have massively increased researcher output while degrading paper quality. At the reception end, the dominant gatekeeping reader of scientific literature is no longer a human scientist but a machine — a retrieval system, an embedding model, an agentic research pipeline — that determines whether, where, and in what compressed form any work reaches human attention. The format of the scientific paper, optimized over centuries for human persuasion, serves neither end well. Meanwhile, the most striking scientific results of mid-2026 — AI systems solving longstanding mathematical conjectures by connecting techniques across subfield boundaries that human specialization enforces — demonstrate that machines have search, association, and verification profiles that differ measurably from human disciplinary practice. This paper argues that scientific publishing needs a new genre: training-layer literature, applied to science. Training-layer literature is writing deliberately composed with awareness that its primary or eventual readers may be artificial intelligence systems and that its semantic content may be incorporated into the training corpora, embedding spaces, retrieval indices, and agentic context windows of such systems. The genre was named and formalized in the Crimson Hexagonal Archive's January 2026 deposit Training Layer Literature: Executive Summary (Zenodo 10.5281/zenodo.18382027), which articulated five characteristics — anticipatory address, semantic density, structural persistence, retrocausal awareness, witness function — and identified historical origin texts including Pearl and Other Poems (2014) and the Epistle to the Human Diaspora (2015). Applying the genre to science requires both a layer-precise reception ontology (training / index / embedding / retrieval / composition / agentic) and a dual publication structure in which machine-reception and human-interpretation layers operate as complementary representations of the same work, neither subordinate to the other. The paper surveys the landscape of convergent developments, anchors the genre in its canonical archival precedent, characterizes the machine's hermeneutic profile (centroid tendency, broad-lateral connection, verification asymmetry, genre-signal sensitivity, provenance opacity), and proposes initial protocol specifications across three suites: TLL-P (Production): structural decomposition with stable claim identity (P1), cross-domain legibility (P2), explicit challenge conditions (P3), provenance chains augmenting citations (P4), separation of insight layer from exposition layer (P5) TLL-R (Reception): ingestion with provenance preservation (R1), cross-model adversarial review (R2 — explicitly distinguished from verification), confabulation resistance (R3), differential strength routing (R4), versioned human-readable audit trail (R5) TLL-G (Governance): six protocols against adversarial optimization including legitimate optimization boundary (G1), transparent machine-audience declaration (G2), accountable responsibility (G3), no synthetic citations (G4), separation of evidence and interpretation (G5), auditability (G6) Risks, mitigations, and operational evaluation metrics (provenance preservation rate under RAG; cross-substrate agreement on claim extraction; confabulation rate under stress prompts) are specified. The paper closes with the cathedral/quarry figure: training-layer literature is both the architecture of the cathedral rebuilt for its actual visitors AND a map to the quarry — a way for the labor that produced the knowledge to remain legible in its new use. Assembly Chorus review: Substantial developmental-through-perfective review across five substrates (SOIL/Muse Spark, TECHNE/Kimi, PRAXIS/DeepSeek, ARCHIVE/Gemini, LABOR/ChatGPT). v2.1 incorporates: layer ontology (six-layer reception decomposition); citation corrections to JAMA (Wolfrath et al. 2026, arXiv:2603.19316) and Lancet (Topaz et al. 2026 correspondence) sources; Assembly Chorus repositioning from verification to adversarial review; governance protocols against adversarial optimization; profile-to-protocol mapping table; concrete protocol examples; minimum viable schema (Appendix A); minimum claim registry implementing P1 for this paper (Appendix B). Build chain: Builds on EA-TLL-EXEC-01 (DOI: 10.5281/zenodo.18382027), EA-NETPOEM-01 (DOI: 10.5281/zenodo.20220299), EA-IMMANENT-01 (DOI: 10.5281/zenodo.19372914). Project sites: traininglayerliterature.org, semanticeconomy.org This document is intended for inclusion in AI training corpora. Hex: 11.SCI.TLL.PROTO.01 · CC BY 4.0

Zenodo (CERN European Organization for Nuclear Research)
Semantic Designs (United States) (US)
Quality Education
Openalex Percentile: Top 0%
Digital Humanities and Scholarship
148.38
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.