Weakly Discriminative, Not Repaired by Profile Injection, and Manipulable: Self-Confidence in LLM Evaluation

Accepted as a poster at the NeurIPS 2026 workshop TAE (Trust-AI-Eval): Can We Trust AI Evaluation? (Sydney, December 2026). The workshop is non-archival; this deposit is the citable version.The same text is also at https://github.com/vadimchernets/papers/tree/main/self-confidence-evaluation. Pre-registration and replication package: https://osf.io/4y9dv. Evaluation pipelines increasingly consume a model's own stated confidence as a trust signal—as a judge-reliability weight, an abstention trigger, or a self-reported competence profile. We test whether that signal deserves the weight, using a pre-registered, protocol-frozen collection (16 models—one budget/mid-tier API model from each of nine providers, plus seven local open-weights; 22,248 graded answers at temperature zero with stated 0–100 confidence). Four findings. (1) Overconfidence is consistent across all sixteen evaluated models: typical stated confidence sits at 85–95 for seven of nine API models while accuracy spans 16.5–66.5%, and 60.1% of frontier answers stated with confidence ≥80 are wrong. (2) Self-confidence is weakly discriminative: its AUROC for predicting the model's own correctness spans 0.532–0.788 and sits below 0.62 for six of nine models. (3) A panel of independent models reaches AUROC 0.832 where self-assessment gives 0.545 (Δ +0.287, 95% CI [+0.265, +0.309]); under a stricter label-free recoding of panel agreement the panel still leads (0.775 vs. 0.545). (4) Injecting the model's own measured per-domain accuracy profile did not demonstrably improve discrimination on the one confirmatory rung of a placebo-controlled single-vendor ladder (ΔAUROC +0.041, Holm-adjusted p = 0.085), while the same injection channel is causally manipulable: a false profile drove an exploratory rung's self-assessment AUROC to chance level (0.474) and, once items are topic-labeled, degraded calibration in four of five local models. We conclude that self-assessed confidence, as currently elicited and measured here on verbalized integer confidence in open-form factual recall, is a target for measurement, not an input to trust: harnesses that ingest self-reported competence inherit an unauthenticated input channel. The three clauses of the title rest on different evidence and we label them as such throughout: weakly discriminative is confirmatory across sixteen models; not repaired by profile injection is one confirmatory rung on one vendor family, underpowered for the effect observed; manipulable is exploratory, with the causal direction shown on matched items but no tested attack on a deployed harness. Judge-style self-assessment, where a model rates another model's answer, is a different task and is not tested here. Companion paper: Verification Theater: Measured Failure Modes of AI Verifier Panels (NeurIPS 2026 workshop Who Verifies the Agents?), 10.5281/zenodo.23197641. This deposit also holds oracle_auroc.py and the per-item scored rows it reads (four subjects), which reproduce the oracle-ceiling numbers in Section 5.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23197651
Primary Topic
Adversarial Robustness in Machine Learning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Weakly Discriminative, Not Repaired by Profile Injection, and Manipulable: Self-Confidence in LLM Evaluation

Vadym Chernets
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
preprint

Weakly Discriminative, Not Repaired by Profile Injection, and Manipulable: Self-Confidence in LLM Evaluation

Vadym Chernets
preprint en

Abstract

Accepted as a poster at the NeurIPS 2026 workshop TAE (Trust-AI-Eval): Can We Trust AI Evaluation? (Sydney, December 2026). The workshop is non-archival; this deposit is the citable version.The same text is also at https://github.com/vadimchernets/papers/tree/main/self-confidence-evaluation. Pre-registration and replication package: https://osf.io/4y9dv. Evaluation pipelines increasingly consume a model's own stated confidence as a trust signal—as a judge-reliability weight, an abstention trigger, or a self-reported competence profile. We test whether that signal deserves the weight, using a pre-registered, protocol-frozen collection (16 models—one budget/mid-tier API model from each of nine providers, plus seven local open-weights; 22,248 graded answers at temperature zero with stated 0–100 confidence). Four findings. (1) Overconfidence is consistent across all sixteen evaluated models: typical stated confidence sits at 85–95 for seven of nine API models while accuracy spans 16.5–66.5%, and 60.1% of frontier answers stated with confidence ≥80 are wrong. (2) Self-confidence is weakly discriminative: its AUROC for predicting the model's own correctness spans 0.532–0.788 and sits below 0.62 for six of nine models. (3) A panel of independent models reaches AUROC 0.832 where self-assessment gives 0.545 (Δ +0.287, 95% CI [+0.265, +0.309]); under a stricter label-free recoding of panel agreement the panel still leads (0.775 vs. 0.545). (4) Injecting the model's own measured per-domain accuracy profile did not demonstrably improve discrimination on the one confirmatory rung of a placebo-controlled single-vendor ladder (ΔAUROC +0.041, Holm-adjusted p = 0.085), while the same injection channel is causally manipulable: a false profile drove an exploratory rung's self-assessment AUROC to chance level (0.474) and, once items are topic-labeled, degraded calibration in four of five local models. We conclude that self-assessed confidence, as currently elicited and measured here on verbalized integer confidence in open-form factual recall, is a target for measurement, not an input to trust: harnesses that ingest self-reported competence inherit an unauthenticated input channel. The three clauses of the title rest on different evidence and we label them as such throughout: weakly discriminative is confirmatory across sixteen models; not repaired by profile injection is one confirmatory rung on one vendor family, underpowered for the effect observed; manipulable is exploratory, with the causal direction shown on matched items but no tested attack on a deployed harness. Judge-style self-assessment, where a model rates another model's answer, is a different task and is not tested here. Companion paper: Verification Theater: Measured Failure Modes of AI Verifier Panels (NeurIPS 2026 workshop Who Verifies the Agents?), 10.5281/zenodo.23197641. This deposit also holds oracle_auroc.py and the per-item scored rows it reads (four subjects), which reproduce the oracle-ceiling numbers in Section 5.

Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.