Algebraic Geometric Empirical Process (AGEP) Theory of Self-Attention Dynamics: From Reduced to Non-Reduced Schemes (v3.11 Unabridged Final Package)
Modern Transformer architectures undergo severe non-equilibrium phase transitions during training, such as representation condensation and rank collapse. In this paper, we establish Algebraic Geometric Empirical Process (AGEP) theory, a unified mathematical paradigm that characterizes Stochastic Gradient Descent (SGD) learning trajectories on singular schemes. First, we formalize a minimal Softmax-free $2 \\times 2$ self-attention mechanism, defining its rank-1 collapse locus as a closed reduced subscheme $S_{\\text{collapse}} = \\text{Spec}(R/\\langle xy-zw \\rangle)$. Applying Hironaka's monoidal blow-up, we prove Theorem 4.1 (Donsker Restoration), demonstrating that the resolution of singularity monomializes the Kullback-Leibler divergence and restores uniform weak convergence (the Donsker property) of the empirical process on exceptional divisor charts, yielding a Real Log Canonical
Authors
- Hideki Ishiyama
Institutions
- Agence France-Presse (FR)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-10
- DOI
- https://doi.org/10.5281/zenodo.22698794
- Primary Topic
- Stochastic Gradient Optimization Techniques
- Type
- preprint