Beyond Accuracy: A Multi-Axis Evaluation Framework for Interpretable Topic Models on Hyperspectral Imagery
Topic models, in particular Latent Dirichlet Allocation (LDA), have been adapted to hyperspectral imagery (HSI) as a route to interpretable spectral mixtures: each spectral unit (canonically a pixel) becomes a document of quantised band-frequency tokens and the inferred topic-word distributions act as non-negative basis spectra. Yet the dominant evaluation criterion remains downstream classification accuracy, an accuracy-only view that is silent on whether the basis is reproducible, band-robust, calibrated, or actually outperforms standard linear baselines. We propose and instantiate a twelve-axis evaluation framework for LDA-style topic models on HSI that quantifies, on the same set of benchmarks: (F-1) classification macro-F1 compared across five method families (paired differences and a hierarchical Bayesian model), (F-2) topic-word coherence (c_v, c_NPMI, U-Mass), (F-3) seed stability for both LDA and two neural variants (ProdLDA, ETM), (F-4) capacity sensitivity across K in {4, 6, 8, 10, 12, 16}, (F-5) band-mask robustness, (F-6) cross-method clustering agreement (LDA vs Felzenszwalb, SLIC, patches, pixels), (F-7) topic-label coupling via P(L | t) entropy and KL, (F-8) per-topic Hungarian-aligned identity tracking, (F-9) HIDSAG cross-preprocessing stability, (F-10) cross-scene topic transfer, (F-11) rate-distortion of the topic representation, and (F-12) external alignment of the topics with a spectral library and a method-rank cross-comparison. We apply the framework to six standard HSI benchmarks (Indian Pines, Salinas, Salinas-A, Pavia University, Kennedy Space Center, Botswana) and five subsets of the HIDSAG mineralogical database (GEOMET, MINERAL1, MINERAL2, GEOCHEM, PORPHYRY). The framework yields three findings that the accuracy-only view does not surface. First, on the labelled scenes the topic-routed soft classifier, which routes the raw spectrum to per-topic logistic experts, matches the raw-spectrum logistic baseline in macro-F1 (0.915 against 0.910, a paired difference of 0.005), while a logistic regression on the topic proportions theta alone trails the baseline by 0.325: the topics add an interpretable routing to the raw spectrum but do not replace it. Second, on Salinas-A under the SWIR-only band mask the paired ARI between canonical and masked dominant-topic maps reaches 0.766, while on Kennedy Space Center under every mask, and on Botswana under the SWIR mask, the same paired ARI collapses to approximately 0.01, separating scenes on which an interpretability claim is band-robust from scenes on which it is not. Third, the theta-only deficit holds on HIDSAG: over its nineteen mineralogical targets the topic-logistic classifier trails raw-logistic by 0.268 macro-F1 (lower on 18 of 19 targets), contrary to the prima-facie expectation that mineralogical mixtures favour topic decomposition over the raw spectrum. We release 3895 deterministic derived artefacts (449 MB), the builder source, the FastAPI backend, the React frontend, and a 109-endpoint test harness as a public, MIT-licensed reproducibility package. Code and derived artefacts: https://github.com/fsantibanezleal/CAOS_LDA_HSI . Interactive web application: https://lda-hsi.fasl-work.com . Manuscript sources: https://github.com/fsantibanezleal/CAOS_LDA_HSI_Paper . Funding: The Advanced Mining Technology Center (AMTC) Basal project (ANID/PIA Project AFB220002) and ANID FONDECYT Postdoctorado 3220094.
Authors
- Felipe Santibañez-Leal (ORCID: https://orcid.org/0000-0002-0150-3246)
Institutions
- Open University of Cyprus (CY)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-19
- DOI
- https://doi.org/10.5281/zenodo.22850118
- Primary Topic
- Remote-Sensing Image Classification
- Type
- preprint