Covariate Copies Count as Evidence in Time-Series Foundation Models

Covariate-aware time-series foundation models forecast a target from a group of related series, and the covariates supplied to them can be redundant: the same sensor exported twice, a temperature in two units, several nearby weather stations. How these models count such redundancy has not been measured against a Bayes reference. We read each model’s implied evidence for a covariate source by finite differences on synthetic worlds in which, for m channels whose errors have correlation ρ, the Bayes-optimal forecast’s log ratio of sensitivities to the two sources changes by log m_eff, with m_eff = m/(1 + (m − 1)ρ). In Chronos-2, t0-beta and TiRex-2, eight exact copies of a covariate raise its implied log evidence ratio by 0.81, 0.68 and 2.28 nats where Bayes predicts zero, 0.33, 0.28 and 1.00 times the response to eight independent measurements, and channels that share part of their measurement error are over-counted as well. On this probe, the number of times a covariate is supplied acts as an implicit, model-specific weight on it. Declaring the source and its ρ inside variate attention and adding − log(m/m_eff) to the logit of each of its keys (a source bias) requires no training. In models whose layers treat every copy of a series alike, it makes the forecast exactly invariant to copies (Proposition 1). On the probe it lowers the worst-case error over all same-source representations to 0.29, 0.27 and 0.69 nats, against at best 0.34, 0.77 and 1.01 for deduplication and preprocessing; paired intervals resolve this difference in t0-beta and TiRex-2 but not in Chronos-2. On GEFCom2012 and GEFCom2017 load forecasting, four copies of the temperature covariate multiply the weighted quantile loss by 0.878 to 1.018 depending on model and dataset, and the source bias returns the forecast made with the temperature supplied once, so that the forecast no longer depends on how many times the covariate was supplied.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23003932
Primary Topic
Meteorological Phenomena and Simulations
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Covariate Copies Count as Evidence in Time-Series Foundation Models

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Meteorological Phenomena and Simulations
preprint

Covariate Copies Count as Evidence in Time-Series Foundation Models

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Covariate-aware time-series foundation models forecast a target from a group of related series, and the covariates supplied to them can be redundant: the same sensor exported twice, a temperature in two units, several nearby weather stations. How these models count such redundancy has not been measured against a Bayes reference. We read each model’s implied evidence for a covariate source by finite differences on synthetic worlds in which, for m channels whose errors have correlation ρ, the Bayes-optimal forecast’s log ratio of sensitivities to the two sources changes by log m_eff, with m_eff = m/(1 + (m − 1)ρ). In Chronos-2, t0-beta and TiRex-2, eight exact copies of a covariate raise its implied log evidence ratio by 0.81, 0.68 and 2.28 nats where Bayes predicts zero, 0.33, 0.28 and 1.00 times the response to eight independent measurements, and channels that share part of their measurement error are over-counted as well. On this probe, the number of times a covariate is supplied acts as an implicit, model-specific weight on it. Declaring the source and its ρ inside variate attention and adding − log(m/m_eff) to the logit of each of its keys (a source bias) requires no training. In models whose layers treat every copy of a series alike, it makes the forecast exactly invariant to copies (Proposition 1). On the probe it lowers the worst-case error over all same-source representations to 0.29, 0.27 and 0.69 nats, against at best 0.34, 0.77 and 1.01 for deduplication and preprocessing; paired intervals resolve this difference in t0-beta and TiRex-2 but not in Chronos-2. On GEFCom2012 and GEFCom2017 load forecasting, four copies of the temperature covariate multiply the weighted quantile loss by 0.878 to 1.018 depending on model and dataset, and the source bias returns the forecast made with the temperature supplied once, so that the forecast no longer depends on how many times the covariate was supplied.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Meteorological Phenomena and Simulations
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.