Specialization Is Not a Myth: Subject-Specific Expert Sets Recur During MoE Generation
Specialization Is Not a Myth: Subject-Specific Expert Sets Recur During MoE Generation. Preprint, version 1.2 (September 2026; version 1.1 August 2026, version 1.0 June 2026). Not peer reviewed. In sparse Mixture-of-Experts models, a learned router selects a small subset of experts for each token from a fixed pool, during both prompt prefill and autoregressive generation. Prior work correctly finds that prompt-side routing can converge on one shared set of experts across domains. We show that generation gives a different answer. We analyze five checkpoints across the Qwen3.5 and GLM-4.7-Flash families, prefill on all five and generation on the two Qwen3.5-35B-A3B checkpoints. In the HauhauCS fine-tune of Qwen3.5-35B-A3B-Instruct, we test whether the selected expert sets follow the local subject of an answer by packing multiple subjects into one generation, letting the model regroup them across prompts, controlling for position, and comparing them with the same questions asked independently. We repeat the subject test on the official instruction-tuned checkpoint from which the HauhauCS model was derived. Position-resolved prefill routing primarily follows the wording of the question, while pooling across the prompt produces the shared usage pattern reported in prior work. Across packed and independently asked versions of identical questions, generation-based expert sets identify the subject at 95 to 96% in Qwen3.5-35B when the answer's register is held constant (0.70 over all design pairs, 0.54 when the register changes), compared with 15 to 17% for prefill and 6% by chance. Generation-time expert sets remain stable within a subject, change sharply at subject boundaries, and recur for the same subject across positions, prompt design, and the official and fine-tuned checkpoints. On the official instruction-tuned checkpoint, generation-time sets pair prompts by subject in 11 of 12 cases where prefill pairs 4 of 12. For short mixed prompts, prefill carries the wording of the question, while generation carries the subject of the answer. Changes in this version. Replicate-then-extend revision with a new title (September 2026). The paper now leads by replicating the shared prefill standing-committee result on five checkpoints across two model families and two gating mechanisms, then shows that generation-time expert sets are organized by the local subject of the answer through a cross-design identification test with wording, position, drift, register, and shuffled-label controls. All version 1.1 results are retained; new results include the identification table, the within-answer boundary analysis, the causal bias test, and the official-checkpoint replication. Archive corrections from the 2026-08-28 audit are disclosed in Methods and Limitations. This deposit contains the paper (LaTeX source and built PDF), the audit trail and reviews (ledger/), the working journal and drafts, the version 1.1 supporting data under data/, and the curated data archive's provenance records, manifests, and complete analysis scripts under curated-data-index/. The full curated archive (about 6 GB of router tensors and per-cell compactions, covered file-by-file by the included MANIFEST.sha256) is available from the author. Files note. This version adds the built PDF (main.pdf) as a standalone file so the paper previews on the record page; the release archive is identical to the 2026-09-01 v1.2 GitHub release.
Authors
- Jeffrey W. Shorthill (ORCID: https://orcid.org/0009-0004-3954-2752)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-01
- DOI
- https://doi.org/10.5281/zenodo.20779604
- Primary Topic
- Expert finding and Q&A systems
- Type
- preprint