Does Graph Structure Earn Its Place in Microservice Root-Cause Analysis? A Controlled Study on RCAEval, and What the Benchmark Was Really Measuring

{"Graph":[0],"neural":[1],"networks":[2],"dominate":[3],"recent":[4,11],"work":[5],"on":[6,44,117,154,217,360],"microservice":[7],"root-cause":[8],"analysis,":[9],"yet":[10],"results":[12,19,405,437],"question":[13],"whether":[14,32],"the":[15,40,67,87,91,107,134,147,151,165,169,202,218,245,285,315,331,335,339,350,371,373,381,390,397,401,436],"graph":[16,68,88,277,351],"contributes.":[17],"Those":[18],"compare":[20],"whole":[21],"pipelines,":[22],"so":[23,356,396],"when":[24],"a":[25,62,139,184,209,222,238,249,254,260,265,269,275,298,304,308,367,414,421,443],"flat":[26,93,299],"model":[27,89,94],"wins":[28],"one":[29,219,410],"cannot":[30],"tell":[31],"structure":[33,357],"is":[34,137,168,183,366,392],"useless":[35],"or":[36],"redundant.":[37],"We":[38,207,418],"run":[39],"comparison":[41],"they":[42],"imply":[43],"RCAEval:":[45],"three":[46,408],"learned":[47],"arms":[48],"with":[49,221,420],"identical":[50],"features,":[51],"optimiser,":[52],"validation":[53],"split,":[54],"early-stopping":[55],"rule":[56],"and":[57,77,133,173,242,274,279,301,349,379,383,413],"scoring":[58,159,240],"head,":[59],"in":[60,131,150,338],"which":[61,196],"single":[63,415],"neighbour-mixing":[64],"term":[65,282],"separates":[66,264],"arms.":[69],"Across":[70],"two":[71,74,110],"RCAEval":[72,119],"benchmarks,":[73],"topology":[75],"sources":[76],"four":[78],"regimes":[79],"we":[80,243],"find":[81],"no":[82,142,387],"reliable":[83],"graph-specific":[84],"effect:":[85],"in-distribution":[86,171,336],"leads":[90],"conventional":[92],"by":[95,442],"0.003":[96],"Avg@5":[97,160,293],"(p":[98],"=":[99,102,313,321],"0.844,":[100],"n":[101,320],"6":[103],"disjoint":[104],"folds).":[105],"Auditing":[106],"pipeline":[108],"surfaced":[109],"benchmark":[111,203],"properties":[112],"that":[113,188],"condition":[114],"any":[115,199],"result":[116],"it.":[118],"injects":[120],"faults":[121],"into":[122,268],"only":[123],"five":[124,153],"services":[125],"per":[126],"system":[127,220,270],"while":[128,434],"exposing":[129],"12-70":[130],"telemetry,":[132],"headline":[135],"metric":[136],"Avg@5:":[138],"ranker":[140],"reading":[141],"telemetry":[143,191,272],"at":[144,319],"all":[145],"places":[146],"true":[148],"culprit":[149],"top":[152],"99.7%":[155],"of":[156,201,427],"held-out":[157],"incidents,":[158],"0.488.":[161],"That":[162],"prior,":[163,271],"not":[164],"uniform-random":[166],"0.137,":[167],"honest":[170],"floor,":[172],"it":[174,225,290,325,445],"collapses":[175],"to":[176,230,347],"0.192":[177],"across":[178],"systems.":[179],"The":[180,257,363],"second":[181],"property":[182],"non-uniform":[185],"column":[186],"schema":[187,224,250],"silently":[189],"zeroes":[190],"for":[192,297,303,330,424],"most":[193],"RE1":[194],"cases,":[195],"changes":[197],"how":[198],"reimplementation":[200],"should":[204],"be":[205],"read.":[206],"reproduce":[208,400],"published":[210],"baseline":[211],"(BARO),":[212],"RCAEval's":[213],"own":[214,232],"reference":[215],"implementation;":[216],"clean":[223],"reaches":[226,291,326,354],"similar":[227],"aggregate":[228],"accuracy":[229],"our":[231],"\\"naive\\"":[233],"heuristic,":[234],"within":[235],"0.004,":[236],"under":[237],"different":[239],"rule,":[241],"read":[244,386],"divergence":[246],"elsewhere":[247],"as":[248],"effect":[251],"rather":[252,369],"than":[253,370],"quality":[255],"difference.":[256],"audit":[258],"motivated":[259,441],"new":[261],"model.":[262],"PSC-GRCA":[263],"candidate":[266],"score":[267],"evidence,":[273],"centred":[276],"residual,":[278],"scores":[280,343],"each":[281,439],"separately.":[283],"On":[284],"six":[286],"fixed":[287],"stratified":[288],"folds":[289],"mean":[292],"0.915":[294],"against":[295,328],"0.864":[296],"MLP":[300],"0.862":[302],"capacity-matched":[305],"no-neighbour":[306],"control,":[307],"fold-level":[309],"paired":[310],"Wilcoxon":[311],"p":[312],"0.03125,":[314],"exact":[316],"two-sided":[317],"minimum":[318],"6.":[322],"Under":[323],"transfer":[324],"0.747":[327],"0.671":[329],"MLP.":[332],"Ablations":[333],"locate":[334],"gain":[337],"prior":[340,388,416],"term:":[341],"prior-only":[342],"0.488,":[344],"prior-free":[345],"drops":[346],"0.850,":[348],"residual":[352,384],"alone":[353],"0.800,":[355],"adds":[358],"little":[359],"its":[361],"own.":[362],"prior-swap":[364],"penalty":[365,391],"guard":[368],"mechanism:":[372],"swap":[374],"passes":[375],"are":[376,406],"made":[377],"deterministic,":[378],"because":[380],"evidence":[382],"branches":[385],"features":[389],"then":[393],"exactly":[394],"zero,":[395],"no-swap":[398],"runs":[399],"swapped":[402],"ones.":[403],"These":[404],"exploratory:":[407],"systems,":[409],"architecture":[411],"family,":[412],"feature.":[417],"close":[419],"twelve-item":[422],"checklist":[423],"this":[425],"class":[426],"study,":[428],"distilled":[429],"from":[430],"sixty-two":[431],"defects":[432],"recorded":[433],"producing":[435],"above,":[438],"item":[440],"failure":[444],"would":[446],"have":[447],"caught.":[448]}

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.22832167
Primary Topic
Software System Performance and Reliability
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Does Graph Structure Earn Its Place in Microservice Root-Cause Analysis? A Controlled Study on RCAEval, and What the Benchmark Was Really Measuring

Imad Buljić
Zenodo (CERN European Organization for Nuclear Research)
Software System Performance and Reliability
preprint

Does Graph Structure Earn Its Place in Microservice Root-Cause Analysis? A Controlled Study on RCAEval, and What the Benchmark Was Really Measuring

Imad Buljić
preprint en

Abstract

Graph neural networks dominate recent work on microservice root-cause analysis, yet recent results question whether the graph contributes. Those results compare whole pipelines, so when a flat model wins one cannot tell whether structure is useless or redundant. We run the comparison they imply on RCAEval: three learned arms with identical features, optimiser, validation split, early-stopping rule and scoring head, in which a single neighbour-mixing term separates the graph arms. Across two RCAEval benchmarks, two topology sources and four regimes we find no reliable graph-specific effect: in-distribution the graph model leads the conventional flat model by 0.003 Avg@5 (p = 0.844, n = 6 disjoint folds). Auditing the pipeline surfaced two benchmark properties that condition any result on it. RCAEval injects faults into only five services per system while exposing 12-70 in telemetry, and the headline metric is Avg@5: a ranker reading no telemetry at all places the true culprit in the top five on 99.7% of held-out incidents, scoring Avg@5 0.488. That prior, not the uniform-random 0.137, is the honest in-distribution floor, and it collapses to 0.192 across systems. The second property is a non-uniform column schema that silently zeroes telemetry for most RE1 cases, which changes how any reimplementation of the benchmark should be read. We reproduce a published baseline (BARO), RCAEval's own reference implementation; on the one system with a clean schema it reaches similar aggregate accuracy to our own "naive" heuristic, within 0.004, under a different scoring rule, and we read the divergence elsewhere as a schema effect rather than a quality difference. The audit motivated a new model. PSC-GRCA separates a candidate score into a system prior, telemetry evidence, and a centred graph residual, and scores each term separately. On the six fixed stratified folds it reaches mean Avg@5 0.915 against 0.864 for a flat MLP and 0.862 for a capacity-matched no-neighbour control, a fold-level paired Wilcoxon p = 0.03125, the exact two-sided minimum at n = 6. Under transfer it reaches 0.747 against 0.671 for the MLP. Ablations locate the in-distribution gain in the prior term: prior-only scores 0.488, prior-free drops to 0.850, and the graph residual alone reaches 0.800, so structure adds little on its own. The prior-swap penalty is a guard rather than the mechanism: the swap passes are made deterministic, and because the evidence and residual branches read no prior features the penalty is then exactly zero, so the no-swap runs reproduce the swapped ones. These results are exploratory: three systems, one architecture family, and a single prior feature. We close with a twelve-item checklist for this class of study, distilled from sixty-two defects recorded while producing the results above, each item motivated by a failure it would have caught.

Zenodo (CERN European Organization for Nuclear Research)
University of Zenica (BA)
Software System Performance and Reliability
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.