Prefailure Transformer Representation Transitions: A Falsifiable Hypothesis for Mathematical Change Before Commitment and Observable Factual Failure

Pre-Failure Transformer Representation Transitions investigates whether an AI model begins to go wrong internally before the failure becomes visible in its final answer. We developed a mathematical framework for measuring changes in the model’s internal representations and identified a measurable transition point where those representations begin to drift from their stable behavior. The work establishes that visible failure may be preceded by an observable internal warning signal, creating the foundation for detecting and potentially intervening in AI failure before the model produces the wrong output.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23068495
Primary Topic
Adversarial Robustness in Machine Learning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Prefailure Transformer Representation Transitions: A Falsifiable Hypothesis for Mathematical Change Before Commitment and Observable Factual Failure

Peregrine Research, Alina Brouwer Guerra, Rivero Carmen E.
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
preprint

Prefailure Transformer Representation Transitions: A Falsifiable Hypothesis for Mathematical Change Before Commitment and Observable Factual Failure

Peregrine Research, Alina Brouwer Guerra, Rivero Carmen E.
preprint en

Abstract

Pre-Failure Transformer Representation Transitions investigates whether an AI model begins to go wrong internally before the failure becomes visible in its final answer. We developed a mathematical framework for measuring changes in the model’s internal representations and identified a measurable transition point where those representations begin to drift from their stable behavior. The work establishes that visible failure may be preceded by an observable internal warning signal, creating the foundation for detecting and potentially intervening in AI failure before the model produces the wrong output.

Zenodo (CERN European Organization for Nuclear Research)
Peregrine Power (United States) (US), Minnesota Project (US)
Peace, Justice and strong institutions
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.