Phase-Native Language Models: A Complex-Valued Alternative to Softmax Attention in Low-Data Regimes
We present Cascade, a fully complex-valued language model that replaces softmaxattention with two lightweight, phase-native operations: learned positional interference andsequential phase modulation. Each token is represented on the complex plane as C[v] = rv ·eiθv ,where the radius rv and angle θv are learned. Interference multiplies each position by aper-position complex scalar; modulation then rotates each position by an angle proportionalto the real part of the previous position’s embedding, providing content-aware messagepassing without queries, keys, or softmax. On a corpus of 8.4M characters from ten public-domain books, and at a matched budget of approximately 500K real-valued parameters,Cascade achieves a held-out perplexity of 5.53, improving on a same-size real-valued MLPby 13.1%, on a single-head attention model by 53.2%, on a two-layer transformer by 19.7%,and on a plain complex phase network by 12.1%. The advantage is stable across randomseeds (+23.5% ± 0.32% at 43K parameters) and grows with model size, from +3.6% at 20Kparameters to +14.2% at 85K on Shakespeare. We ablate eleven phase-native mechanismsand find that only the specific combination of interference and modulation yields substantialgains, while pure phase embeddings alone match real-valued embeddings. We release thecomplete codebase, all training scripts, and raw experimental logs.
Authors
- Masoud Azizi (ORCID: https://orcid.org/0000-0002-8864-0533)
Institutions
- Freelancer (Portugal) (PT)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23044272
- Primary Topic
- Topic Modeling
- Type
- preprint