Post-hoc Interpretability of LDA on Hyperspectral Imagery: SHAP Attributions, Counterfactual Topic Flips, and LLM-judge Alignment under Token-Mass-Dispersion Asymmetry

{"The":[0,449],"interpretability":[1,41,416],"claim":[2],"of":[3,31,60,125,178,285,300,346],"LDA-on-HSI":[4],"rests":[5],"on":[6,43],"the":[7,19,22,29,32,44,53,136,172,186,193,208,250,297,301,369,373,383,396,402,414,420],"assumption":[8],"that":[9,104],"topic-word":[10],"distributions":[11],"are":[12,105,337],"human-readable.":[13],"This":[14],"study":[15],"is":[16,28,188,192,198,211,232,291],"restricted":[17],"to":[18,72,112,117,123,320,375],"LDA":[20],"backbone;":[21],"cross-backbone":[23],"comparison":[24,370],"(HDP,":[25],"ProdLDA,":[26],"ETM)":[27],"subject":[30],"companion":[33,54],"paper":[34,55],"P4.":[35],"We":[36,93,406],"test":[37],"three":[38,95,274],"orthogonal":[39],"post-hoc":[40],"axes":[42],"V1-V15,":[45],"V17-V20":[46],"wordification":[47],"sweep":[48],"(nineteen":[49],"LDA-fitted":[50],"recipes)":[51],"from":[52],"P3:":[56],"F-13":[57,99,409],"SHAP":[58,100,110,410],"attribution":[59],"pixel-to-topic":[61],"decisions":[62],"via":[63],"a":[64,74,83,152,162,175,204,246,281,292],"closed-form":[65],"posterior,":[66],"F-22":[67,129,418],"counterfactual":[68,130],"L1":[69,131],"perturbation":[70],"required":[71],"flip":[73,181],"document's":[75,84,363],"argmax":[76,89],"topic,":[77],"and":[78,87,121,161,229,238,252,267,313,332,352,391,423,434,460],"F-15":[79,378,424],"LLM-as-judge":[80],"alignment":[81],"between":[82,227],"top":[85,91],"tokens":[86,277,287],"its":[88,215,286,289,335,357,365,427],"topic's":[90,247],"tokens.":[92],"find":[94],"concrete":[96],"results.":[97],"First,":[98],"gives":[101],"recipe-specific":[102],"explanations":[103],"consistent":[106],"across":[107,322],"scenes:":[108],"V1's":[109],"reduces":[111],"specific":[113,118],"wavelength":[114],"bands,":[115,367],"V7's":[116],"absorption":[119],"features,":[120],"V12's":[122],"clusters":[124],"Gaussian-mixture":[126],"components.":[127],"Second,":[128],"(sentinel-patched":[132],"19-recipe":[133],"ranking)":[134],"separates":[135],"recipes":[137,272],"into":[138],"an":[139,392],"\\"ultra-robust\\"":[140],"band":[141,154,164,174],"(V20,":[142],"V12,":[143],"V3;":[144],"sentinel-patched":[145],"means":[146],"26.3":[147],"/":[148,150],"24.5":[149],"23.5),":[151],"\\"moderate\\"":[153],"(V1":[155],"=":[156,159,166,169,344],"6.1,":[157],"V7":[158],"5.2)":[160],"\\"fragile\\"":[163],"(V9":[165],"1.0,":[167],"V10":[168],"1.2);":[170],"within":[171,182],"ultra-robust":[173],"large":[176],"fraction":[177],"documents":[179,222,228],"never":[180,265],"50":[183],"steps,":[184],"so":[185,245,294,368],"ordering":[187],"not":[189],"strict":[190],"(V12":[191],"most":[194,241],"robust":[195],"where":[196],"flip-sampling":[197],"adequate).":[199],"Third,":[200],"F-15,":[201],"computed":[202],"with":[203,273,426],"deterministic":[205],"stand-in":[206],"for":[207,236,253,271,350,354],"LLM":[209,393],"judge,":[210],"governed":[212],"by":[213,225,234],"how":[214],"top-10":[216,248,290,364],"overlap":[217],"rule":[218,302],"reads":[219],"each":[220,284,362],"recipe's":[221],"rather":[223,307],"than":[224,308,388],"agreement":[226],"topics.":[230],"It":[231],"1.0":[233,270],"construction":[235,404,428],"V2":[237],"V8":[239],"(at":[240],"12":[242],"word":[243],"types,":[244],"covers":[249],"vocabulary)":[251],"V9":[254],"(a":[255],"one-token":[256],"document":[257,282],"can":[258],"be":[259],"judged":[260],"aligned":[261],"or":[262,275],"ambiguous":[263],"but":[264,356],"misaligned),":[266],"stays":[268],"near":[269],"four":[276],"per":[278],"document.":[279],"Where":[280],"holds":[283],"once,":[288],"tie,":[293],"we":[295],"report":[296],"exact":[298],"expectation":[299],"under":[303],"uniformly":[304],"random":[305],"tie-breaking":[306],"one":[309],"sort":[310],"order:":[311],"V3":[312,351],"V12":[314],"then":[315],"score":[316],"0.17,":[317],"against":[318,348],"0.03":[319],"0.18":[321],"individual":[323],"orders.":[324],"V20":[325],"re-weights":[326],"V3's":[327],"identical":[328],"(band,":[329],"bin)":[330],"alphabet":[331],"scores":[333],"0.65;":[334],"topics":[336],"also":[338,360],"less":[339],"dispersed":[340],"(effective":[341],"vocabulary":[342],"N_eff":[343],"exp(H(phi_k))":[345],"309":[347],"431":[349],"400":[353],"V12),":[355],"per-band":[358],"copies":[359],"make":[361],"highest-copy":[366],"cannot":[371],"attribute":[372],"gap":[374],"topic":[376,389],"dispersion.":[377],"as":[379,413,419],"defined":[380],"therefore":[381],"measures":[382],"documents'":[384],"token":[385,398],"structure":[386],"more":[387],"interpretability,":[390],"oracle":[394],"given":[395],"same":[397,403],"lists":[399],"would":[400],"face":[401],"effects.":[405],"recommend":[407],"reporting":[408],"top-K":[411],"attributions":[412],"primary":[415],"artefact,":[417,422],"topic-stability":[421],"only":[425],"effects":[429],"reported":[430],"beside":[431],"it.":[432],"Code":[433],"derived":[435],"artefacts:":[436],"https://github.com/fsantibanezleal/CAOS_LDA_HSI":[437],".":[438,443,447],"Interactive":[439],"web":[440],"application:":[441],"https://lda-hsi.fasl-work.com":[442],"Manuscript":[444],"sources:":[445],"https://github.com/fsantibanezleal/CAOS_LDA_HSI_Paper":[446],"Funding:":[448],"Advanced":[450],"Mining":[451],"Technology":[452],"Center":[453],"(AMTC)":[454],"Basal":[455],"project":[456],"(ANID/PIA":[457],"Project":[458],"AFB220002)":[459],"ANID":[461],"FONDECYT":[462],"Postdoctorado":[463],"3220094.":[464]}

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.21504113
Citations
25
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Post-hoc Interpretability of LDA on Hyperspectral Imagery: SHAP Attributions, Counterfactual Topic Flips, and LLM-judge Alignment under Token-Mass-Dispersion Asymmetry

Felipe Santibañez-Leal
25 citations
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Post-hoc Interpretability of LDA on Hyperspectral Imagery: SHAP Attributions, Counterfactual Topic Flips, and LLM-judge Alignment under Token-Mass-Dispersion Asymmetry

Felipe Santibañez-Leal
preprint en
25 citations

Abstract

The interpretability claim of LDA-on-HSI rests on the assumption that topic-word distributions are human-readable. This study is restricted to the LDA backbone; the cross-backbone comparison (HDP, ProdLDA, ETM) is the subject of the companion paper P4. We test three orthogonal post-hoc interpretability axes on the V1-V15, V17-V20 wordification sweep (nineteen LDA-fitted recipes) from the companion paper P3: F-13 SHAP attribution of pixel-to-topic decisions via a closed-form posterior, F-22 counterfactual L1 perturbation required to flip a document's argmax topic, and F-15 LLM-as-judge alignment between a document's top tokens and its argmax topic's top tokens. We find three concrete results. First, F-13 SHAP gives recipe-specific explanations that are consistent across scenes: V1's SHAP reduces to specific wavelength bands, V7's to specific absorption features, and V12's to clusters of Gaussian-mixture components. Second, F-22 counterfactual L1 (sentinel-patched 19-recipe ranking) separates the recipes into an "ultra-robust" band (V20, V12, V3; sentinel-patched means 26.3 / 24.5 / 23.5), a "moderate" band (V1 = 6.1, V7 = 5.2) and a "fragile" band (V9 = 1.0, V10 = 1.2); within the ultra-robust band a large fraction of documents never flip within 50 steps, so the ordering is not strict (V12 is the most robust where flip-sampling is adequate). Third, F-15, computed with a deterministic stand-in for the LLM judge, is governed by how its top-10 overlap rule reads each recipe's documents rather than by agreement between documents and topics. It is 1.0 by construction for V2 and V8 (at most 12 word types, so a topic's top-10 covers the vocabulary) and for V9 (a one-token document can be judged aligned or ambiguous but never misaligned), and stays near 1.0 for recipes with three or four tokens per document. Where a document holds each of its tokens once, its top-10 is a tie, so we report the exact expectation of the rule under uniformly random tie-breaking rather than one sort order: V3 and V12 then score 0.17, against 0.03 to 0.18 across individual orders. V20 re-weights V3's identical (band, bin) alphabet and scores 0.65; its topics are also less dispersed (effective vocabulary N_eff = exp(H(phi_k)) of 309 against 431 for V3 and 400 for V12), but its per-band copies also make each document's top-10 its highest-copy bands, so the comparison cannot attribute the gap to topic dispersion. F-15 as defined therefore measures the documents' token structure more than topic interpretability, and an LLM oracle given the same token lists would face the same construction effects. We recommend reporting F-13 SHAP top-K attributions as the primary interpretability artefact, F-22 as the topic-stability artefact, and F-15 only with its construction effects reported beside it. Code and derived artefacts: https://github.com/fsantibanezleal/CAOS_LDA_HSI . Interactive web application: https://lda-hsi.fasl-work.com . Manuscript sources: https://github.com/fsantibanezleal/CAOS_LDA_HSI_Paper . Funding: The Advanced Mining Technology Center (AMTC) Basal project (ANID/PIA Project AFB220002) and ANID FONDECYT Postdoctorado 3220094.

Zenodo (CERN European Organization for Nuclear Research)
Open University of Cyprus (CY)
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.