Learning Predictive Memory: Adaptation and Length Extrapolation in Time-Series Transformers
Modern time-series architectures encode temporal dependence through very different mechanisms, including sparse retrieval, exponential recurrence, heavy-tailed attention, and seasonal memory. We study these choices through a common object, the predictive-memory kernel that governs one-step forecasting. This turns architecture design into approximation and estimation in a process-dependent forecast-risk geometry. We characterize the statistical complexity of several structured temporal memories and construct a same-realization forecaster that adapts among sparse, exponential, power-law, and seasonal classes. We then use Hankel structure to quantify the forecasting cost of using finite-state memory for an incompatible temporal law. We finally study context-length extrapolation. Changing the softmax support rescales the realized predictive-memory coefficients even when the relative temporal law is itself correct at the longer context, the mechanism by which length-dependent softmax dispersion degrades forecasting. We derive the excess-risk floor this produces, compare it with the truncated-truth oracle, and obtain a memory-law-exact normalization correction. For genuinely content-dependent attention, the correction becomes sample-dependent. Controlled experiments confirm the statistical and normalization predictions. A frozen-model intervention on trained Transformers moves held-out error in the predicted direction at every context tested and significantly so on average across them.Code reproducing all figures is included in this record.
Authors
- Yuheng Song
Institutions
- Huawei Technologies (China) (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22757506
- Primary Topic
- Forecasting Techniques and Applications
- Type
- preprint