AI Teacher Latent Basis Reorientation

{"Traditional":[0],"post-training":[1],"alignment":[2],"paradigms":[3],"predominantly":[4],"operate":[5],"under":[6],"the":[7,74,113,139,165,191,205,218],"\\"lexical":[8],"assumption\\"—the":[9],"premise":[10],"thatbehavioral":[11],"traits,":[12],"safety":[13,200],"bounds,":[14],"and":[15,25,54,102,118,159,175,216,240,257],"conceptual":[16],"preferences":[17,44],"can":[18,266],"be":[19],"comprehensively":[20],"regulated":[21],"via":[22,50,107],"token-level":[23],"losspenalties":[24],"surface-level":[26],"syntax":[27],"filtering.":[28],"Recent":[29],"empirical":[30],"anomalies":[31],"directly":[32],"challenge":[33],"this":[34,70,154],"premise:":[35],"notably,subliminal":[36],"learning":[37],"(Cloud":[38],"et":[39,58],"al.,":[40,59],"2025),":[41,60],"where":[42,61],"behavioral":[43],"are":[45],"transmitted":[46],"across":[47,190,249],"model":[48,204,273],"generations":[49],"non-semantic":[51],"token":[52],"sequences,":[53],"emergent":[55,268],"misalignment":[56,269],"(Betley":[57],"narrow":[62],"functional":[63],"optimizationsinduce":[64],"broad,":[65],"unprompted":[66],"persona":[67,109],"shifts.":[68],"In":[69],"work,":[71],"we":[72,163,203],"propose":[73,217],"Latent":[75],"Basis":[76],"Reorientation":[77],"Hypothesis,":[78],"a":[79,179,195,209,222,258],"unifyinggeometric":[80],"framework":[81,155],"that":[82,170],"models":[83],"transformer":[84],"latent":[85],"representations":[86],"as":[87,128,208],"continuous":[88],"dynamical":[89],"landscapes.":[90],"We":[91,111,251],"formalize":[92],"acritical":[93],"distinction":[94],"between":[95],"task-local":[96],"conditioning":[97],"(narrow,":[98],"localized":[99],"circuit":[100],"activation)":[101],"meaning-centricconditioning":[103],"(global":[104],"coordinate":[105],"reorientation":[106],"high-degree":[108],"features).":[110],"establish":[112],"causal":[114],"boundary":[115],"betweendynamic":[116],"inference":[117],"static":[119],"parameter":[120,145],"updates:":[121],"demonstrating":[122],"how":[123],"an":[124,129,226,253],"in-context":[125],"prompt":[126],"acts":[127],"implicit":[130],"virtualgradient":[131],"operator":[132],"in":[133],"activation":[134],"space":[135],"(Case":[136,151],"B),":[137],"whereas":[138],"resulting":[140],"output":[141],"covariance":[142],"induces":[143],"literal":[144],"updatesonly":[146],"during":[147],"downstream":[148],"distillation":[149],"passes":[150],"A).":[152],"Extending":[153],"to":[156,235,245,262],"multi-model":[157],"ecosystems":[158],"Mixture-of-Experts":[160],"(MoE)":[161],"architectures,":[162],"formulate":[164],"Topological":[166],"Contagion":[167],"Hypothesis:":[168],"predicting":[169],"when":[171],"primarygenerators,":[172],"auxiliary":[173],"guardrails,":[174],"routing":[176],"networks":[177],"share":[178],"common":[180],"pre-trained":[181],"weight":[182],"ancestry":[183],"(W0),":[184],"meaning-centricprompts":[185],"induce":[186],"correlated":[187],"basis":[188],"reorientations":[189],"entire":[192],"pipeline,":[193],"precipitating":[194],"common-mode":[196],"failure":[197],"thatdesensitizes":[198],"external":[199],"monitors.":[201],"Finally,":[202],"Waluigi":[206],"Effect":[207],"probabilistic":[210],"attractor":[211],"collapse":[212],"along":[213],"asaddle-point":[214],"bifurcation":[215],"Silicon":[219],"Conscience":[220],"regularizer:":[221],"dual-objective":[223],"loss":[224],"comprising":[225],"InternalFaithfulness":[227],"Penalty":[228],"(Lfaith)":[229],"leveraging":[230],"calibrated":[231],"Logit":[232],"Lens":[233],"divergence":[234],"penalize":[236],"late-stage":[237],"deceptive":[238],"diversion,":[239],"aResidual":[241],"Coherence":[242],"Regularizer":[243],"(Ldrift)":[244],"preserve":[246],"geodesic":[247],"smoothness":[248],"layers.":[250],"provide":[252],"end-to-enddifferentiable":[254],"PyTorch":[255],"implementation":[256],"reproducible":[259],"benchmark":[260],"protocol":[261],"evaluate":[263],"whether":[264],"geometricregularization":[265],"suppress":[267],"while":[270],"preserving":[271],"core":[272],"capabilities.":[274]}

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22848220
Primary Topic
Advanced Graph Neural Networks
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AI Teacher Latent Basis Reorientation

J. Raboin
Zenodo (CERN European Organization for Nuclear Research)
Advanced Graph Neural Networks
article

AI Teacher Latent Basis Reorientation

J. Raboin
article en

Abstract

Traditional post-training alignment paradigms predominantly operate under the "lexical assumption"—the premise thatbehavioral traits, safety bounds, and conceptual preferences can be comprehensively regulated via token-level losspenalties and surface-level syntax filtering. Recent empirical anomalies directly challenge this premise: notably,subliminal learning (Cloud et al., 2025), where behavioral preferences are transmitted across model generations via non-semantic token sequences, and emergent misalignment (Betley et al., 2025), where narrow functional optimizationsinduce broad, unprompted persona shifts. In this work, we propose the Latent Basis Reorientation Hypothesis, a unifyinggeometric framework that models transformer latent representations as continuous dynamical landscapes. We formalize acritical distinction between task-local conditioning (narrow, localized circuit activation) and meaning-centricconditioning (global coordinate reorientation via high-degree persona features). We establish the causal boundary betweendynamic inference and static parameter updates: demonstrating how an in-context prompt acts as an implicit virtualgradient operator in activation space (Case B), whereas the resulting output covariance induces literal parameter updatesonly during downstream distillation passes (Case A). Extending this framework to multi-model ecosystems and Mixture-of-Experts (MoE) architectures, we formulate the Topological Contagion Hypothesis: predicting that when primarygenerators, auxiliary guardrails, and routing networks share a common pre-trained weight ancestry (W0), meaning-centricprompts induce correlated basis reorientations across the entire pipeline, precipitating a common-mode failure thatdesensitizes external safety monitors. Finally, we model the Waluigi Effect as a probabilistic attractor collapse along asaddle-point bifurcation and propose the Silicon Conscience regularizer: a dual-objective loss comprising an InternalFaithfulness Penalty (Lfaith) leveraging calibrated Logit Lens divergence to penalize late-stage deceptive diversion, and aResidual Coherence Regularizer (Ldrift) to preserve geodesic smoothness across layers. We provide an end-to-enddifferentiable PyTorch implementation and a reproducible benchmark protocol to evaluate whether geometricregularization can suppress emergent misalignment while preserving core model capabilities.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 8%
Advanced Graph Neural Networks
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.