Topo-DiT: Balanced Routing Bottlenecks for Latent Video Flow Matching
Long video sequences make dense transformer computation expensive. We investigate a latent-grid flow model that aggregates spatiotemporal features into a smaller set of feature slots using balanced entropic transport, processes these slots with a time-conditioned transformer, and broadcasts the resulting features back to the latent grid. The proposed contribution is a testable routing design, not a claim that compact tokens inherently represent physical objects. We specify consistent transport marginals, latent-space training and sampling, an optional masked reconstruction warmup, and matched routing ablations. The implementation passes 26 CPU-only correctness tests without video-model training. No video-generation quality, physical reasoning, memory advantage, or distributed-scaling result is claimed. The current artifact is intended to support subsequent controlled experiments under limited local compute. code: https://github.com/shandingwangyue/topo_dit
Authors
- lipeng (ORCID: https://orcid.org/0009-0001-7774-4142)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22955031
- Primary Topic
- Video Coding and Compression Technologies
- Type
- article
- Field-Weighted Citation Impact
- 0.00