Grokking the Discrete Logarithm: Scaling Across Model Capacity and Mechanistic Analysis
To the best of our knowledge, we provide the first systematic demonstration that a trans-former can exhibit grokking on the full variable-base finite-field discrete logarithm mapping,(g, h) → x with g^x ≡ h (mod p). A 427k-parameter, two-layer transformer reaches 99.8% test accuracy at p = 113 only after a long memorization phase, giving a clear instance of delayed algorithmic generalization on a structured inverse problem. We then perform a 12-model capacity sweep spanning 90k–14.4M parameters. Grokking is observed at the427k anchor and again from 1.64M through 14.4M parameters under a staged training-dataschedule, with a non-monotonic region whose outcome changes when training data areincreased. Mechanistically, the grokking transition is accompanied by increasing alignmentwith multiplicative-character structure. Symbolic regression on hidden states preferentiallyrecovers character-aligned trigonometric features, while a free operator set independentlyrecovers Fourier harmonics. Finally, prime-order and disjoint-generator controls show that the phenomenon is not restricted to smooth composite-order groups or to generators seen during training. These results position DLP as a compact benchmark for studying howmemorization, representation formation, and generalization interact across neural scale
Authors
- David Vesterlund
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-17
- DOI
- https://doi.org/10.5281/zenodo.22813397
- Primary Topic
- Ferroelectric and Negative Capacitance Devices
- Type
- preprint