Entangled Alignment: When Safety Is the Substrate
Post-training alignment has substantially improved behavior, but whether safety-relevant habits of judgment would be more durable if learned during capability formation remains unknown. Entangled Alignment proposes practicing understanding and care together within the same pretraining examples from which capabilities are learned. Pretraining corpora preserve the library more fully than the reader: what authors wrote is recorded more often than the questions, connections, evaluations, and belief revisions through which understanding changes. We call this omission the Missing Reader. A Teacher constructs chronological, source-adjacent reader traces; a Student would learn from them during pretraining. The full treatment gives one continuing Reader a provenance-linked graph for retaining and revising understanding across sources, and a compact first-person Reader Core recited in full at every thinking boundary—a policy we call Total Saturation. Context modulates the evaluation that follows, never the recital. Trace, graph, and Core are separately testable: the question is whether the learned orientation becomes consequential and survives continued training, targeted erosion, and successor-like transfer without semantic drift. Two Teacher-side prototype case studies produced inspectable graph and trace artifacts that were then audited. No Student has yet been trained; these studies do not show that the traces reproduce expert-quality reading or that the graph or Core causes any benefit. Entangled Alignment is therefore a falsifiable research program, not a demonstrated alignment method. Its motivating question is whether this training can move a Student from merely becoming the text toward becoming the reader.
Authors
- Henrik Westerberg (ORCID: https://orcid.org/0000-0002-7204-6900)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22739447
- Primary Topic
- Text Readability and Simplification
- Type
- preprint