Measurability Before Power: Pre-Confirmatory Viability Screening for Sparse LLM Experiments

Experiments comparing large language model (LLM) interventions on sparse software-engineering tasks can spend most of their budget before establishing that a comparison is measurable. We adapt established multi-endpoint feasibility methodology to sparse LLM author/reviewer experiments, operationalise three viability endpoints, and report a prospectively governed case in which the screen stopped a confirmatory study before its reserved evaluation pool was consumed. Two development pilots, each with 1,320 model calls and 840 judged outcomes, failed criteria fixed before model calls. Task-level outcomes moved on 4 of 40 and 0 of 40 tasks. A replication-aware movement statistic has a positive, success-rate-dependent reference null even under zero treatment effect; evaluated at each pilot's pooled baseline, observed responsiveness was 0.36 and 0.00 times the pooled-binomial independent-arm reference null. Reviewers used a structural NO ISSUES FOUND channel on 0.0–8.3% of drafts against a prospective 10% trigger, while the planning model was near-separated or failed to fit. The prospective stop rests on the frozen endpoints, not on the reference-null comparison or the post-hoc estimability diagnostic. External construct-alignment tests and assumption-dependent operating-characteristic simulations are supporting evidence only. The external analysis supplies no positive corroboration after a post-hoc, reviewer-motivated noise null. We contribute an LLM-specific operationalisation of established feasibility methodology and a documented case in which a screen stopped a study before its reserved evidence was spent; we claim no new statistical method, causal result about critique, or external validation.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22810166
Primary Topic
Software Engineering Research
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Measurability Before Power: Pre-Confirmatory Viability Screening for Sparse LLM Experiments

Sai Varun Thupakula
Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
preprint

Measurability Before Power: Pre-Confirmatory Viability Screening for Sparse LLM Experiments

Sai Varun Thupakula
preprint en

Abstract

Experiments comparing large language model (LLM) interventions on sparse software-engineering tasks can spend most of their budget before establishing that a comparison is measurable. We adapt established multi-endpoint feasibility methodology to sparse LLM author/reviewer experiments, operationalise three viability endpoints, and report a prospectively governed case in which the screen stopped a confirmatory study before its reserved evaluation pool was consumed. Two development pilots, each with 1,320 model calls and 840 judged outcomes, failed criteria fixed before model calls. Task-level outcomes moved on 4 of 40 and 0 of 40 tasks. A replication-aware movement statistic has a positive, success-rate-dependent reference null even under zero treatment effect; evaluated at each pilot's pooled baseline, observed responsiveness was 0.36 and 0.00 times the pooled-binomial independent-arm reference null. Reviewers used a structural NO ISSUES FOUND channel on 0.0–8.3% of drafts against a prospective 10% trigger, while the planning model was near-separated or failed to fit. The prospective stop rests on the frozen endpoints, not on the reference-null comparison or the post-hoc estimability diagnostic. External construct-alignment tests and assumption-dependent operating-characteristic simulations are supporting evidence only. The external analysis supplies no positive corroboration after a post-hoc, reviewer-motivated noise null. We contribute an LLM-specific operationalisation of established feasibility methodology and a documented case in which a screen stopped a study before its reserved evidence was spent; we claim no new statistical method, causal result about critique, or external validation.

Zenodo (CERN European Organization for Nuclear Research)
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Measurability Before Power: Pre-Confirmatory Viability Screening for Sparse LLM Experiments — Sai Varun Thupakula · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS