TransLearn: Representing and Comparing Transformer Architectures in a Free Categorical Framework
We define TransLearn, a free category of residual sequential architecture terms generated by a typed primitive signature. Its base doctrine is Cartesian with specified commutative-monoid operations on residual objects, rather than biproducts, so nonlinear Euclidean layers remain admissible. For a fixed signature, finite-graph completeness identifies the terms with finite acyclic typed architecture graphs modulo the stated Cartesian, monoid, and sequence-coherence equations; pure signature extensions are conservative on old terms. The syntax separates parameter identity, runtime routing, causality, state, and implementation labels. Finite depth-sharing patterns form the partition lattice; nested routed families flatten to a single routed family when dispatch and combine morphisms are unrestricted; and support certificates give a sufficient prefix-causality test for length-changing residual branches. Applied to Funnel-Transformer's published stride-2 mean pooling and repetition upsampling, the test leads to an explicit counterexample showing that their direct composition is noncausal if reused unchanged inside an autoregressive residual branch. The theory decides these structural questions; it does not decide optimization, accuracy, or hardware efficiency without additional interpretations or cost assumptions.
Authors
- Alexander A
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-06
- DOI
- https://doi.org/10.5281/zenodo.22550792
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- preprint