Rank-Adaptive Local Empirical Processes in Low-Rank Attention
We study the non-asymptotic statistical complexity of rank-restricted linear projection layersin multi-head self-attention mechanisms under localized quadratic loss. Recent empirical studiesobserve that attention projection matrices often exhibit severe rank collapse during training,concentrating parameter updates on low-rank subvarieties M≤s. To analyze generalization underthis geometric restriction, we establish a distribution-uniform local empirical process frameworkfor linear projection matrices fW (X) = W X. Localizing the loss class within an excess-risk ballK(W) ≤ r, we derive a distribution-uniform L2(Q)-bracketing metric entropy bound of orderO(s(dout +din)log ((1 + 8M_R√s ρ_r)/ε),where s denotes the matrix rank bound. Applying Dudley’sentropy integral to the localized envelope yields a non-asymptotic local Rademacher complexityshrinkage rate of O(((√{s(dout+din)r_n})/( √{λmin n}))(√{1 + log(1 + s)})) as excess risk rn ↓ 0. Furthermore, solving the sub-root fixed point equation yields an excess-risk fast estimation rate of O((s(dout+din) log s)/(λmin n)).For a 4096×4096 projection layer restricted to rank s = 16, this bound reflects a 99.22% reductionin manifold degrees of freedom (peff = 130, 816 vs. 1.68 × 107) and a corresponding contractionin local Rademacher complexity relative to full-rank limits. We also establish the fixed-dimensionP-Donsker property for these localized classes and derive an output prediction-distortion boundunder truncated SVD projections.
Authors
- Hideki Ishiyama
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23147548
- Primary Topic
- Stochastic Gradient Optimization Techniques
- Type
- preprint