Prefailure Transformer Representation Transitions: A Falsifiable Hypothesis for Mathematical Change Before Commitment and Observable Factual Failure
Pre-Failure Transformer Representation Transitions investigates whether an AI model begins to go wrong internally before the failure becomes visible in its final answer. We developed a mathematical framework for measuring changes in the model’s internal representations and identified a measurable transition point where those representations begin to drift from their stable behavior. The work establishes that visible failure may be preceded by an observable internal warning signal, creating the foundation for detecting and potentially intervening in AI failure before the model produces the wrong output.
Authors
- Peregrine Research
- Alina Brouwer Guerra
- Rivero Carmen E.
Institutions
- Peregrine Power (United States) (US)
- Minnesota Project (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23068495
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- preprint