Can Algorithm-Aligned Input Representations Support Latent Algorithm Learning? A Case Study Using Multi-Digit Multiplication
As demonstrated by prior work on algorithmic reasoning, standard Transformers can struggle to learn reliable arithmetic from input–output examples. Existing approaches provide support through scratchpads, chain-of-thought intermediate supervision, positional representations, or arithmetic-specific modules with built-in operations. This paper explores relational input structure: what if we just give a standard Transformer a better input representation containing the proper algorithm-aligned inductive bias? I test this idea using multi-digit multiplication as a controlled case study. A learned lattice frontend organizes every pair of operand digits before passing the resulting representation to an otherwise standard Transformer encoder–decoder. The grid organization is fixed, while the cell features are learned; no local products, carries, diagonal sums, scratchpads, or intermediate labels are supplied. With a validation-gated curriculum that gradually increases operand length while replaying earlier lengths, the resulting 11.8-million-parameter model achieves 99.775% aggregate exact-match accuracy across 20,000 newly generated equal-length operand pairs spanning one through 20 digits, including 96.5% accuracy at 20 digits. PCA and decoder cross-attention reveal multiplication-relevant internal organization. Together, these results provide strong empirical support for latent algorithm learning within the trained length range and motivate further study of algorithm-aligned input representations.
Authors
- David He
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23195779
- Primary Topic
- Topic Modeling
- Type
- preprint