Let Token Draw Itself

Recent autoregressive image generation models achieve impressive results by modeling images as sequences of discrete tokens. However, these tokens are treated as opaque latent codes that are decoded jointly, making the visual contribution of each generation step invisible. We propose that each token should instead independently carry its own visual content as a self-contained drawing primitive. We present TokenPainter, where every token independently produces an RGB color and spatial mask via per-token heads, composited via softmax into the final image. Without auxiliary losses, softmax competition drives tokens to spontaneously specialize into complementary spatial regions, from broad silhouettes to fine details. To train TokenPainter, we introduce GroupPos, a GAN loss that normalizes discriminator scores within each group via z-score and applies a soft non-negative threshold. TokenPainter trained with a fixed number of tokens extrapolates to different token counts at inference while preserving subject identity, and distinct subsets of tokens recombine into structurally consistent outputs, demonstrating token-level modularity.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23149095
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Let Token Draw Itself

Yujing Tang
Zenodo (CERN European Organization for Nuclear Research)
Generative Adversarial Networks and Image Synthesis
preprint

Let Token Draw Itself

Yujing Tang
preprint en

Abstract

Recent autoregressive image generation models achieve impressive results by modeling images as sequences of discrete tokens. However, these tokens are treated as opaque latent codes that are decoded jointly, making the visual contribution of each generation step invisible. We propose that each token should instead independently carry its own visual content as a self-contained drawing primitive. We present TokenPainter, where every token independently produces an RGB color and spatial mask via per-token heads, composited via softmax into the final image. Without auxiliary losses, softmax competition drives tokens to spontaneously specialize into complementary spatial regions, from broad silhouettes to fine details. To train TokenPainter, we introduce GroupPos, a GAN loss that normalizes discriminator scores within each group via z-score and applies a soft non-negative threshold. TokenPainter trained with a fixed number of tokens extrapolates to different token counts at inference while preserving subject identity, and distinct subsets of tokens recombine into structurally consistent outputs, demonstrating token-level modularity.

Zenodo (CERN European Organization for Nuclear Research)
University of Chinese Academy of Sciences (CN)
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Let Token Draw Itself — Yujing Tang · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS