Which Wordification Matters? A Nineteen-Recipe Sweep of the Interpretable-Topic-Model Framework on Hyperspectral Imagery
The companion paper (P1) introduces a twelve-axis evaluation framework for topic models on hyperspectral imagery and instantiates it on a single canonical wordification recipe (V1, band-frequency tokenisation). This paper asks a different question: how does the choice of wordification itself shape the conclusions of that framework. We define a family of nineteen wordification recipes V1-V15, V17-V20 spanning seven axes of design freedom (token alphabet, spatial vs. spectral aggregation, local vs. global vocabulary, document-length regime, signal transform, label-aware weighting, and learnt representation) and sweep them through the framework on the six labelled-scene panel (Indian Pines, Salinas and Salinas-A, Pavia University, Kennedy Space Center and Botswana). The headline finding is that there is no universal winner: V1 wins F-2 coherence on 2/6 scenes; V3, a joint (band, bin) Cartesian vocabulary, wins F-7 topic-label normalised MI on 2/6 scenes; V12 (Gaussian-mixture tokens) wins F-2 on 2/6 and F-7 on 2/6; V20 (mutual-information-weighted bands, whose weights come from the labels F-7 is scored against) wins F-2 coherence (0.88) and F-7 NMI (0.44) together on Indian Pines; F-1 macro-F1 separates the recipes by less than 0.01 in six-scene mean (0.912 to 0.922; paired over scenes and folds, V12 leads every other recipe by 0.004 to 0.009), and V20's F-1 is affected by label leakage through its MI weighting, as caveated in the F-1 protocol. V20 has the highest counterfactual robustness F-22 of the sweep at Q=8 (L1 = 26.3 vs 24.5 for V12), a token count that grows with its documents, about six times longer than V3's. The full 19-recipe Q-sweep at Q in {8, 16, 32} identifies three recipes whose F-7 NMI and F-2 c_v both improve monotonically: V20, V2 (intensity-bin), and V8 (NFINDR endmember). V20 has the largest F-2 gain of the three (+0.060) and gains +0.043 on F-7 (V2: +0.045); it reaches the highest absolute F-7 value at Q=32 (0.563); on F-2 c_v it ties V12 at Q=32 (0.910 vs 0.909, margin 0.0016, a 3-3 per-scene split). At Q=8 V20 trails V12 on the LDA F-7 mean by 0.014 NMI; the ranking inverts at Q=32 where V20 leads the LDA F-7 mean by 0.030 and outscores V12 on 5/6 scenes (it wins F-7 outright on 4/6: Salinas-A, Pavia U, KSC and Botswana). Of the other fifteen recipes, seven decline monotonically on F-7 and four peak at Q=16 then regress. A four-backbone x nineteen-recipe extension of F-7 NMI identifies V8 (NFINDR endmember-fraction) as the most cross-backbone-consistent recipe by a clear margin (mean 0.431); V20 (0.397) and V2 (0.395) rank second and third, with V20 the only label-aware recipe in the consistently-strong family. The per-backbone F-7 winners are V12 (LDA, 0.534), V11 (HDP, 0.571), V8 (ProdLDA, 0.328) and V20 (ETM, 0.490). V13 (vector-quantised VAE) underperforms sharply; V18 (graph-Laplacian eigenvector tokens) places third on F-7 mean among the extension recipes (V13-V20). The per-axis winner depends both on the scene and on which property of the topic basis the axis measures. We report the nineteen-recipe result matrix on the six labelled scenes for F-1, F-2 and F-7, with the extension axes F-13, F-14, F-17, F-18 and F-22, identify four axis-recipe affinities, and recommend that V1 retain its canonical status as a reproducibility default rather than as a universal best. V12, V3 and V20 have the three highest six-scene means on both F-2 coherence and F-7 normalised MI, and on F-7 they lie within 0.014 NMI of each other, suggesting that the design-space ceiling for fixed-K LDA on the six labelled scenes is close to saturation; a foundation-model wordification (V16) is left to a separate study. Code and derived artefacts: https://github.com/fsantibanezleal/CAOS_LDA_HSI . Interactive web application: https://lda-hsi.fasl-work.com . Manuscript sources: https://github.com/fsantibanezleal/CAOS_LDA_HSI_Paper . Funding: The Advanced Mining Technology Center (AMTC) Basal project (ANID/PIA Project AFB220002) and ANID FONDECYT Postdoctorado 3220094.
Authors
- Felipe Santibañez-Leal (ORCID: https://orcid.org/0000-0002-0150-3246)
Institutions
- Open University of Cyprus (CY)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.21504117
- Citations
- 25
- Primary Topic
- Geochemistry and Geologic Mapping
- Type
- preprint