PRIMITIVE COMBINATORIAL ARCHITECTURE OF DNA

PRIMITIVE COMBINATORIAL ARCHITECTURE OF DNA A Probabilistic Model of Patterns, Cycles, and Structural Differentiation from Four Symbols Author: Cláudio Vicente da SilvaAffiliation: Independent Researcher — Londrina, Paraná, BrazilDate: September 17, 2026 Description This research develops a formal mathematical and conceptual architecture for representing DNA as a finite combinatorial system generated from four fundamental symbols: A={A,T,C,G},∣A∣=4.\\mathcal{A}=\\{A,T,C,G\\},\\qquad |\\mathcal{A}|=4. For sequences of length nn, the corresponding combinatorial space is defined as: Sn=An,∣Sn∣=4n.\\mathcal{S}_n=\\mathcal{A}^n,\\qquad |\\mathcal{S}_n|=4^n. The study begins with the combinatorial structure of the four-symbol alphabet and progressively develops a formal framework for sequence generation, probability, composition, complementarity, replication, differentiation, Hamming distance, state transitions, deterministic cycles, and accessible versus potential sequence spaces. The central architectural development introduces a distinction between Latency, Potency, Act, and Transduction. In formal terms, the architecture is represented as: Lt→Pt→At→Tt→Lt+1\\boxed{ \\mathcal{L}_t\\rightarrow\\mathcal{P}_t\\rightarrow\\mathcal{A}_t\\rightarrow\\mathcal{T}_t\\rightarrow\\mathcal{L}_{t+1} } and, in its integrated form: LATENCY→POTENCY→ACT→TRANSDUCTION→NEW STATE\\boxed{ \\text{LATENCY}\\rightarrow \\text{POTENCY}\\rightarrow \\text{ACT}\\rightarrow \\text{TRANSDUCTION}\\rightarrow \\text{NEW STATE} } Here, latency represents a space of possible configurations; potency represents conditions governing realization; act represents an instantiated state; and transduction represents the transformation connecting one state to another. These terms constitute structural components of the proposed architecture and are not asserted to be numerically identical quantities. The research further establishes a molecular representation: VC→XD→TIXT\\boxed{ V_C\\rightarrow X_D\\xrightarrow{T_I}X_T } where VCV_C denotes a documented causal genetic condition or variant, XDX_D the associated biological state, TIT_I the intervention or transformation, and XTX_T the subsequent observed state. A significant structural criterion of the model is that the causal genetic location does not necessarily have to coincide with the therapeutic intervention location: LC≠LI\\boxed{ L_C\\neq L_I } This permits the same architecture to represent different intervention modalities, including gene addition, viral-vector delivery, ex vivo cellular transduction, and genome editing. External Biological Cases The proposed representation was subsequently examined against five independently documented therapeutic contexts: OTOF and genetic hearing loss, involving gene replacement through an AAV-based therapeutic approach. WAS and Wiskott–Aldrich syndrome, involving ex vivo modification of autologous hematopoietic stem and progenitor cells. SMN1 and spinal muscular atrophy, involving delivery of a functional SMN sequence. RPE65 and inherited retinal dystrophy, involving AAV-mediated gene addition. HBB-related sickle cell disease and Casgevy, in which the causal genetic condition is associated with HBB while the therapeutic intervention acts at the erythroid BCL11A regulatory system. These cases were not treated as proofs of a universal biological law. Their role is to test whether a common formal representation can accommodate documented biological systems with different genes, variants, tissues, intervention mechanisms, molecular locations, and measurable outcomes. The study therefore distinguishes explicitly between: HYPOTHESIS≠MATHEMATICAL CONSTRUCTION≠OBSERVED DATA≠RESULT≠INTERPRETATION\\boxed{ \\text{HYPOTHESIS} \\neq \\text{MATHEMATICAL CONSTRUCTION} \\neq \\text{OBSERVED DATA} \\neq \\text{RESULT} \\neq \\text{INTERPRETATION} } Formal Validation Bank The final stage of the research transforms the initial case-based analysis into a general validation protocol. Each biological case is represented as: Mi=(Gi,Vi,Li,Xi,Ti,LTi,Ri)\\boxed{ M_i=(G_i,V_i,L_i,X_i,T_i,L_{T_i},R_i) } and the complete validation bank as: B={M1,M2,…,MN}.\\boxed{ \\mathcal{B}=\\{M_1,M_2,\\ldots,M_N\\}. } The protocol defines explicit criteria for: completeness of documentary information; representation of a biological case by the architecture; partial representation; structural incompatibility; model modification; exception recording; independence of evidence units; diversity of genes, variants, biological states, interventions, locations, and outcomes; generalization to cases not used in the initial construction. The central quantitative measure is the representation rate: ρ=NRN\\boxed{ \\rho=\\frac{N_R}{N} } where NRN_R is the number of represented cases and NN is the total number of evaluated cases. For sufficiently complete cases, an additional measure is defined: ρcomplete=NRN−NP\\boxed{ \\rho_{\\mathrm{complete}} = \\frac{N_R}{N-N_P} } where NPN_P represents partially documented cases. A separate modification index records how frequently the original architecture requires structural alteration: A=Ncases requiring modificationNcomplete cases\\boxed{ A= \\frac{N_{\\mathrm{cases\\ requiring\\ modification}}} {N_{\\mathrm{complete\\ cases}}} } These quantities are methodological measures of representational performance within a defined validation sample. They are not presented as measures of universal biological truth. Generalization and Falsification The proposed protocol establishes a progression from an initial five-case representation toward progressively larger validation banks: 5→10→25→50→100→N.5\\rightarrow10\\rightarrow25\\rightarrow50\\rightarrow100\\rightarrow N. For each expansion, the representation rate can be recalculated: ρN=NRN.\\boxed{ \\rho_N=\\frac{N_R}{N}. } The architecture is required to remain fixed during a predefined validation round. A new case should not automatically generate a new rule. If previously undocumented cases repeatedly require new structural rules, such modifications must be explicitly recorded. The generalization test is formulated as: B=Btrain∪Btest\\boxed{ \\mathcal{B} = \\mathcal{B}_{\\mathrm{train}} \\cup \\mathcal{B}_{\\mathrm{test}} } with the test cases evaluated without changing the fundamental rules established before their classification. The central methodological question is therefore: HOW MANY REAL CASES CAN THE ARCHITECTURE REPRESENT WITHOUT MODIFICATION?\\boxed{ \\text{HOW MANY REAL CASES CAN THE ARCHITECTURE REPRESENT WITHOUT MODIFICATION?} } Scope of the Contribution The work combines discrete combinatorics, probability, sequence representation, state-transition systems, molecular mapping, and a formal validation protocol. Its original mathematical and architectural components include: the four-symbol combinatorial representation of DNA; sequence-space construction; probability and composition models; complementarity as an involutive operation; replication and preservation models; Hamming-distance analysis; deterministic and probabilistic transition structures; finite-state cycle representation; accessible versus potential sequence spaces; the Latency–Potency–Act–Transduction architecture; molecular mapping between causal state, intervention, and subsequent state; a formal validation-bank structure; representation and completeness metrics; exception and modification criteria; independent-evidence accounting; sample-diversity documentation; train/test generalization procedures; pre-specification of validation criteria. The research does not claim that the proposed architecture has been established as a universal biological law. Rather, it establishes a formal framework that can be subjected to progressively larger documentary and computational tests. The methodological progression is summarized as: CONSTRUCTION→APPLICATION→EXTERNAL VALIDATION→FORMAL VALIDATION BANK→QUANTITATIVE TEST\\boxed{ \\text{CONSTRUCTION} \\rightarrow \\text{APPLICATION} \\rightarrow \\text{EXTERNAL VALIDATION} \\rightarrow \\text{FORMAL VALIDATION BANK} \\rightarrow \\text{QUANTITATIVE TEST} } The research consequently proposes a reproducible path from a mathematical construction to external testing and quantitative assessment. Data and Evidence Status The biological cases incorporated into the validation stage are based on publicly documented scientific, genomic, clinical, and regulatory sources. These sources provide the external biological data used for molecular and therapeutic mapping. The external sources do not constitute validation of the mathematical architecture by themselves. Instead, they provide independent cases against which the architecture can be evaluated. The study therefore maintains a strict separation between source data and the author's formal construction. Research Status This work is presented as an independent research contribution and methodological framework. The validation bank is designed to remain open to expansion, replication, independent auditing, additional biological cases, and possible identification of structural incompatibilities. The principal research objective of the subsequent validation stage is not to presuppose confirmation, but to measure the representational capacity of the architecture under predefined criteria. Central methodological principle: ONE MODEL→MULTIPLE CASES\\boxed{ \\text{ONE MODEL}\\rightarrow\\text{MULTIPLE CASES} } rather than constructing a separate theoretical structure for every individual case.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22818739
Primary Topic
DNA and Nucleic Acid Chemistry
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

PRIMITIVE COMBINATORIAL ARCHITECTURE OF DNA

Vicente da Silva
Zenodo (CERN European Organization for Nuclear Research)
DNA and Nucleic Acid Chemistry
preprint

PRIMITIVE COMBINATORIAL ARCHITECTURE OF DNA

Vicente da Silva
preprint en

Abstract

PRIMITIVE COMBINATORIAL ARCHITECTURE OF DNA A Probabilistic Model of Patterns, Cycles, and Structural Differentiation from Four Symbols Author: Cláudio Vicente da SilvaAffiliation: Independent Researcher — Londrina, Paraná, BrazilDate: September 17, 2026 Description This research develops a formal mathematical and conceptual architecture for representing DNA as a finite combinatorial system generated from four fundamental symbols: A={A,T,C,G},∣A∣=4.\mathcal{A}=\{A,T,C,G\},\qquad |\mathcal{A}|=4. For sequences of length nn, the corresponding combinatorial space is defined as: Sn=An,∣Sn∣=4n.\mathcal{S}_n=\mathcal{A}^n,\qquad |\mathcal{S}_n|=4^n. The study begins with the combinatorial structure of the four-symbol alphabet and progressively develops a formal framework for sequence generation, probability, composition, complementarity, replication, differentiation, Hamming distance, state transitions, deterministic cycles, and accessible versus potential sequence spaces. The central architectural development introduces a distinction between Latency, Potency, Act, and Transduction. In formal terms, the architecture is represented as: Lt→Pt→At→Tt→Lt+1\boxed{ \mathcal{L}_t\rightarrow\mathcal{P}_t\rightarrow\mathcal{A}_t\rightarrow\mathcal{T}_t\rightarrow\mathcal{L}_{t+1} } and, in its integrated form: LATENCY→POTENCY→ACT→TRANSDUCTION→NEW STATE\boxed{ \text{LATENCY}\rightarrow \text{POTENCY}\rightarrow \text{ACT}\rightarrow \text{TRANSDUCTION}\rightarrow \text{NEW STATE} } Here, latency represents a space of possible configurations; potency represents conditions governing realization; act represents an instantiated state; and transduction represents the transformation connecting one state to another. These terms constitute structural components of the proposed architecture and are not asserted to be numerically identical quantities. The research further establishes a molecular representation: VC→XD→TIXT\boxed{ V_C\rightarrow X_D\xrightarrow{T_I}X_T } where VCV_C denotes a documented causal genetic condition or variant, XDX_D the associated biological state, TIT_I the intervention or transformation, and XTX_T the subsequent observed state. A significant structural criterion of the model is that the causal genetic location does not necessarily have to coincide with the therapeutic intervention location: LC≠LI\boxed{ L_C\neq L_I } This permits the same architecture to represent different intervention modalities, including gene addition, viral-vector delivery, ex vivo cellular transduction, and genome editing. External Biological Cases The proposed representation was subsequently examined against five independently documented therapeutic contexts: OTOF and genetic hearing loss, involving gene replacement through an AAV-based therapeutic approach. WAS and Wiskott–Aldrich syndrome, involving ex vivo modification of autologous hematopoietic stem and progenitor cells. SMN1 and spinal muscular atrophy, involving delivery of a functional SMN sequence. RPE65 and inherited retinal dystrophy, involving AAV-mediated gene addition. HBB-related sickle cell disease and Casgevy, in which the causal genetic condition is associated with HBB while the therapeutic intervention acts at the erythroid BCL11A regulatory system. These cases were not treated as proofs of a universal biological law. Their role is to test whether a common formal representation can accommodate documented biological systems with different genes, variants, tissues, intervention mechanisms, molecular locations, and measurable outcomes. The study therefore distinguishes explicitly between: HYPOTHESIS≠MATHEMATICAL CONSTRUCTION≠OBSERVED DATA≠RESULT≠INTERPRETATION\boxed{ \text{HYPOTHESIS} \neq \text{MATHEMATICAL CONSTRUCTION} \neq \text{OBSERVED DATA} \neq \text{RESULT} \neq \text{INTERPRETATION} } Formal Validation Bank The final stage of the research transforms the initial case-based analysis into a general validation protocol. Each biological case is represented as: Mi=(Gi,Vi,Li,Xi,Ti,LTi,Ri)\boxed{ M_i=(G_i,V_i,L_i,X_i,T_i,L_{T_i},R_i) } and the complete validation bank as: B={M1,M2,…,MN}.\boxed{ \mathcal{B}=\{M_1,M_2,\ldots,M_N\}. } The protocol defines explicit criteria for: completeness of documentary information; representation of a biological case by the architecture; partial representation; structural incompatibility; model modification; exception recording; independence of evidence units; diversity of genes, variants, biological states, interventions, locations, and outcomes; generalization to cases not used in the initial construction. The central quantitative measure is the representation rate: ρ=NRN\boxed{ \rho=\frac{N_R}{N} } where NRN_R is the number of represented cases and NN is the total number of evaluated cases. For sufficiently complete cases, an additional measure is defined: ρcomplete=NRN−NP\boxed{ \rho_{\mathrm{complete}} = \frac{N_R}{N-N_P} } where NPN_P represents partially documented cases. A separate modification index records how frequently the original architecture requires structural alteration: A=Ncases requiring modificationNcomplete cases\boxed{ A= \frac{N_{\mathrm{cases\ requiring\ modification}}} {N_{\mathrm{complete\ cases}}} } These quantities are methodological measures of representational performance within a defined validation sample. They are not presented as measures of universal biological truth. Generalization and Falsification The proposed protocol establishes a progression from an initial five-case representation toward progressively larger validation banks: 5→10→25→50→100→N.5\rightarrow10\rightarrow25\rightarrow50\rightarrow100\rightarrow N. For each expansion, the representation rate can be recalculated: ρN=NRN.\boxed{ \rho_N=\frac{N_R}{N}. } The architecture is required to remain fixed during a predefined validation round. A new case should not automatically generate a new rule. If previously undocumented cases repeatedly require new structural rules, such modifications must be explicitly recorded. The generalization test is formulated as: B=Btrain∪Btest\boxed{ \mathcal{B} = \mathcal{B}_{\mathrm{train}} \cup \mathcal{B}_{\mathrm{test}} } with the test cases evaluated without changing the fundamental rules established before their classification. The central methodological question is therefore: HOW MANY REAL CASES CAN THE ARCHITECTURE REPRESENT WITHOUT MODIFICATION?\boxed{ \text{HOW MANY REAL CASES CAN THE ARCHITECTURE REPRESENT WITHOUT MODIFICATION?} } Scope of the Contribution The work combines discrete combinatorics, probability, sequence representation, state-transition systems, molecular mapping, and a formal validation protocol. Its original mathematical and architectural components include: the four-symbol combinatorial representation of DNA; sequence-space construction; probability and composition models; complementarity as an involutive operation; replication and preservation models; Hamming-distance analysis; deterministic and probabilistic transition structures; finite-state cycle representation; accessible versus potential sequence spaces; the Latency–Potency–Act–Transduction architecture; molecular mapping between causal state, intervention, and subsequent state; a formal validation-bank structure; representation and completeness metrics; exception and modification criteria; independent-evidence accounting; sample-diversity documentation; train/test generalization procedures; pre-specification of validation criteria. The research does not claim that the proposed architecture has been established as a universal biological law. Rather, it establishes a formal framework that can be subjected to progressively larger documentary and computational tests. The methodological progression is summarized as: CONSTRUCTION→APPLICATION→EXTERNAL VALIDATION→FORMAL VALIDATION BANK→QUANTITATIVE TEST\boxed{ \text{CONSTRUCTION} \rightarrow \text{APPLICATION} \rightarrow \text{EXTERNAL VALIDATION} \rightarrow \text{FORMAL VALIDATION BANK} \rightarrow \text{QUANTITATIVE TEST} } The research consequently proposes a reproducible path from a mathematical construction to external testing and quantitative assessment. Data and Evidence Status The biological cases incorporated into the validation stage are based on publicly documented scientific, genomic, clinical, and regulatory sources. These sources provide the external biological data used for molecular and therapeutic mapping. The external sources do not constitute validation of the mathematical architecture by themselves. Instead, they provide independent cases against which the architecture can be evaluated. The study therefore maintains a strict separation between source data and the author's formal construction. Research Status This work is presented as an independent research contribution and methodological framework. The validation bank is designed to remain open to expansion, replication, independent auditing, additional biological cases, and possible identification of structural incompatibilities. The principal research objective of the subsequent validation stage is not to presuppose confirmation, but to measure the representational capacity of the architecture under predefined criteria. Central methodological principle: ONE MODEL→MULTIPLE CASES\boxed{ \text{ONE MODEL}\rightarrow\text{MULTIPLE CASES} } rather than constructing a separate theoretical structure for every individual case.

Zenodo (CERN European Organization for Nuclear Research)
Sustainable cities and communities
DNA and Nucleic Acid Chemistry
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.