Which Backbone Picks Which Wordification? A Factorial Study of Topic-Model Families on Hyperspectral Imagery

The companion paper (P3) showed that nineteen wordification recipes (V1-V15, V17-V20) produce measurably different topic bases under a fixed LDA backbone, with no universal winner: V12 (Gaussian-mixture tokens) has the highest mean coherence and topic-label coupling (winning two of six scenes on each), V20 (mutual-information-weighted bands) has the highest counterfactual robustness F-22, and V1 (band-frequency tokens, the canonical) is the most reliable of the per-band recipes under reseeding, while F-1 macro-F1 separates the recipes by less than 0.01. This paper asks the orthogonal question: does the recipe winner depend on the topic-model backbone? We run the full V1-V20 sweep (nineteen recipes; V16 is scaffolded only) under three additional backbones, namely the Hierarchical Dirichlet Process (HDP, nonparametric K), the amortised neural topic model labelled ProdLDA (run in its mixture-decoder form, with a standard logistic-normal prior) and the Embedded Topic Model (ETM, the same prior and encoder with a word-topic embedding decoder), producing a 4 x 19 factorial on six labelled-scene benchmarks (456 cells in the headline F-2 coherence axis). The headline finding is that the wordification ranking inverts across backbones. LDA and ETM agree that V12 wins, HDP picks V7 (absorption-feature triplets) and ProdLDA picks V3 (joint (band, bin) vocabulary). ETM and ProdLDA differ only in the decoder parameterisation, so their different winners come from the decoder or from optimisation; the runs do not separate the two. We give a per-recipe affinity score that quantifies how each recipe transfers across backbones, identify three empirical backbone-recipe pairings, and recommend reporting the (backbone, wordification) pair as a single hyper-decision in any future LDA-on-HSI work. LDVAE-T, a Dirichlet-prior unmixing VAE, is a candidate fifth backbone; it is not evaluated because no public implementation is available. Code and derived artefacts: https://github.com/fsantibanezleal/CAOS_LDA_HSI . Interactive web application: https://lda-hsi.fasl-work.com . Manuscript sources: https://github.com/fsantibanezleal/CAOS_LDA_HSI_Paper . Funding: The Advanced Mining Technology Center (AMTC) Basal project (ANID/PIA Project AFB220002) and ANID FONDECYT Postdoctorado 3220094.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22850111
Primary Topic
Neural dynamics and brain function
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Which Backbone Picks Which Wordification? A Factorial Study of Topic-Model Families on Hyperspectral Imagery

Felipe Santibañez-Leal
Zenodo (CERN European Organization for Nuclear Research)
Neural dynamics and brain function
preprint

Which Backbone Picks Which Wordification? A Factorial Study of Topic-Model Families on Hyperspectral Imagery

Felipe Santibañez-Leal
preprint en

Abstract

The companion paper (P3) showed that nineteen wordification recipes (V1-V15, V17-V20) produce measurably different topic bases under a fixed LDA backbone, with no universal winner: V12 (Gaussian-mixture tokens) has the highest mean coherence and topic-label coupling (winning two of six scenes on each), V20 (mutual-information-weighted bands) has the highest counterfactual robustness F-22, and V1 (band-frequency tokens, the canonical) is the most reliable of the per-band recipes under reseeding, while F-1 macro-F1 separates the recipes by less than 0.01. This paper asks the orthogonal question: does the recipe winner depend on the topic-model backbone? We run the full V1-V20 sweep (nineteen recipes; V16 is scaffolded only) under three additional backbones, namely the Hierarchical Dirichlet Process (HDP, nonparametric K), the amortised neural topic model labelled ProdLDA (run in its mixture-decoder form, with a standard logistic-normal prior) and the Embedded Topic Model (ETM, the same prior and encoder with a word-topic embedding decoder), producing a 4 x 19 factorial on six labelled-scene benchmarks (456 cells in the headline F-2 coherence axis). The headline finding is that the wordification ranking inverts across backbones. LDA and ETM agree that V12 wins, HDP picks V7 (absorption-feature triplets) and ProdLDA picks V3 (joint (band, bin) vocabulary). ETM and ProdLDA differ only in the decoder parameterisation, so their different winners come from the decoder or from optimisation; the runs do not separate the two. We give a per-recipe affinity score that quantifies how each recipe transfers across backbones, identify three empirical backbone-recipe pairings, and recommend reporting the (backbone, wordification) pair as a single hyper-decision in any future LDA-on-HSI work. LDVAE-T, a Dirichlet-prior unmixing VAE, is a candidate fifth backbone; it is not evaluated because no public implementation is available. Code and derived artefacts: https://github.com/fsantibanezleal/CAOS_LDA_HSI . Interactive web application: https://lda-hsi.fasl-work.com . Manuscript sources: https://github.com/fsantibanezleal/CAOS_LDA_HSI_Paper . Funding: The Advanced Mining Technology Center (AMTC) Basal project (ANID/PIA Project AFB220002) and ANID FONDECYT Postdoctorado 3220094.

Zenodo (CERN European Organization for Nuclear Research)
Open University of Cyprus (CY)
Peace, Justice and strong institutions
Neural dynamics and brain function
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.