Let Token Draw Itself
Recent autoregressive image generation models achieve impressive results by modeling images as sequences of discrete tokens. However, these tokens are treated as opaque latent codes that are decoded jointly, making the visual contribution of each generation step invisible. We propose that each token should instead independently carry its own visual content as a self-contained drawing primitive. We present TokenPainter, where every token independently produces an RGB color and spatial mask via per-token heads, composited via softmax into the final image. Without auxiliary losses, softmax competition drives tokens to spontaneously specialize into complementary spatial regions, from broad silhouettes to fine details. To train TokenPainter, we introduce GroupPos, a GAN loss that normalizes discriminator scores within each group via z-score and applies a soft non-negative threshold. TokenPainter trained with a fixed number of tokens extrapolates to different token counts at inference while preserving subject identity, and distinct subsets of tokens recombine into structurally consistent outputs, demonstrating token-level modularity.
Authors
- Yujing Tang
Institutions
- University of Chinese Academy of Sciences (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23149095
- Primary Topic
- Generative Adversarial Networks and Image Synthesis
- Type
- preprint