Algebraic Geometric Empirical Process (AGEP) Theory of Self-Attention Dynamics: From Reduced to Non-Reduced Schemes (v3.11 Unabridged Final Package)

Modern Transformer architectures undergo severe non-equilibrium phase transitions during training, such as representation condensation and rank collapse. In this paper, we establish Algebraic Geometric Empirical Process (AGEP) theory, a unified mathematical paradigm that characterizes Stochastic Gradient Descent (SGD) learning trajectories on singular schemes. First, we formalize a minimal Softmax-free $2 \\times 2$ self-attention mechanism, defining its rank-1 collapse locus as a closed reduced subscheme $S_{\\text{collapse}} = \\text{Spec}(R/\\langle xy-zw \\rangle)$. Applying Hironaka's monoidal blow-up, we prove Theorem 4.1 (Donsker Restoration), demonstrating that the resolution of singularity monomializes the Kullback-Leibler divergence and restores uniform weak convergence (the Donsker property) of the empirical process on exceptional divisor charts, yielding a Real Log Canonical

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-10
DOI
https://doi.org/10.5281/zenodo.22698794
Primary Topic
Stochastic Gradient Optimization Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Algebraic Geometric Empirical Process (AGEP) Theory of Self-Attention Dynamics: From Reduced to Non-Reduced Schemes (v3.11 Unabridged Final Package)

Hideki Ishiyama
Zenodo (CERN European Organization for Nuclear Research)
Stochastic Gradient Optimization Techniques
preprint

Algebraic Geometric Empirical Process (AGEP) Theory of Self-Attention Dynamics: From Reduced to Non-Reduced Schemes (v3.11 Unabridged Final Package)

Hideki Ishiyama
preprint en

Abstract

Modern Transformer architectures undergo severe non-equilibrium phase transitions during training, such as representation condensation and rank collapse. In this paper, we establish Algebraic Geometric Empirical Process (AGEP) theory, a unified mathematical paradigm that characterizes Stochastic Gradient Descent (SGD) learning trajectories on singular schemes. First, we formalize a minimal Softmax-free $2 \times 2$ self-attention mechanism, defining its rank-1 collapse locus as a closed reduced subscheme $S_{\text{collapse}} = \text{Spec}(R/\langle xy-zw \rangle)$. Applying Hironaka's monoidal blow-up, we prove Theorem 4.1 (Donsker Restoration), demonstrating that the resolution of singularity monomializes the Kullback-Leibler divergence and restores uniform weak convergence (the Donsker property) of the empirical process on exceptional divisor charts, yielding a Real Log Canonical

Zenodo (CERN European Organization for Nuclear Research)
Agence France-Presse (FR)
Sustainable cities and communities
Stochastic Gradient Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Algebraic Geometric Empirical Process (AGEP) Theory of Self-Attention Dynamics: From Reduced to Non-Reduced Schemes (v3.11 Unabridged Final Package) — Hideki Ishiyama · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS