Identifiability of Ensemble Disagreement under One-Step Distillation
Model-based reinforcement learning gates imagined transitions by how much an ensemble disagrees, and diffusion world models are distilled into one-step students to make that imagination affordable. Whether distillation preserves the score the gate reads is rarely checked; we show that it need not. Matching the teacher's value-mapped spread from a single shared latent constrains one scalar combination of two distinct quantities—disagreement across ensemble members and conditional spread within a member—so neither component is identified. A student can therefore match the scalar loss while producing a different uncertainty value, including the opposite sign; empirically, this can translate into degraded uncertainty ranking. Two or more shared latents restore identification of both components in population. In one-step distillation of a diffusion ensemble on a delayed-branch control task, the one-latent student improves next-state mean squared error in all 30 seeds while losing teacher–student Spearman correlation of the uncertainty score in all 30 (mean drop 0.31). Inverse-squared-scale reweighting effectively suppresses the within-member term on this map, and equal weights recover most of the lost ranking. The practical picture is mixed: ordinary member matching retains higher uncertainty-ranking agreement on hopper-hop, including under one frozen critic. The identifiability result covers the distilled scalar statistic rather than the full trained loss, and shared-latent agreement is a coupled measure, not the independent-noise agreement a deployed gate would draw
Authors
- keyush nisar
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23167797
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- preprint