Stratum: Lie Field Layers for Grid Logic Puzzles in Small Transformers
We introduce Lie field layers, which add a decaying, optionally rotating recurrence over sequence positions to a transformer’s residual stream, and evaluate them on abstract-vocabulary grid logic puzzles. This version 1.1 corrects errors in version 1.0; every change is listed in Section 9. On 3×3 puzzles, a small transformer with a Lie field layer reaches 0.933 ± 0.022 puzzle accuracy against 0.874 ± 0.037 for the baseline across 13 seeds per arm (p=0.0001 by a rank test with one poorly converged baseline seed excluded, a choice made after the results were seen; p<0.001 with it included). A decay-only ablation with the same parameter count shows that the SO(d) rotation contributes (+0.039, p=0.0075), although its share of the gain is poorly determined; whether decay alone helps is unresolved. On 4×4 puzzles we find no mean improvement: across 15 seeds the layer gives 0.096 ± 0.032 against 0.108 ± 0.014 for the baseline, with large seed-to-seed variation. Removing the sigmoid gate does not collapse accuracy, and we withdraw version 1.0’s claim that the gate is necessary. Validation loss ranks trained models well at the checkpoints we report but is a poor stopping criterion: at 3×3 it rises while accuracy is still improving. The puzzles contain no ordering relations, a limitation of the data generator that we disclose.
Authors
- Curtis Sandoval
Institutions
- Lexmark (United States) (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22848974
- Primary Topic
- VLSI and FPGA Design Techniques
- Type
- preprint