The fools are certain; the wise are doubtful: Exploring LLM confidence in code completion

Abstract Code completion entails the task of providing missing tokens given a surrounding context. It can boost developer productivity and serve as a code discovery tool. Code completion has recently been approached with Large Language Models (LLMs) fine-tuned on code (code LLMs). The performance of code LLMs can be assessed with downstream and intrinsic metrics. Downstream metrics are usually employed to evaluate the practical utility of a model, but can be unreliable and require complex calculations and domain-specific knowledge. In contrast, intrinsic metrics such as perplexity, entropy, and mutual information, which measure model confidence or uncertainty, are simple, versatile, and universal across LLMs and tasks, and have been proposed as proxies for functional correctness and hallucination risk in LLM-generated code. Motivated by this, we evaluate the confidence of LLMs when generating code by measuring code perplexity across programming languages, models, and datasets using various LLMs, and a sample of 2254 files from 881 GitHub projects. We find that strongly-typed languages exhibit lower perplexity than dynamically typed languages. Scripting languages also demonstrate higher perplexity. Shell appears universally high in perplexity, whereas Java appears low. Code perplexity depends on the employed LLM; under a fixed model, relative language-level rankings are largely stable across evaluation corpora. Although code comments modestly increase perplexity, the language ranking based on perplexity is barely affected by their presence. LLM researchers, developers, and users can use our findings to assess the suitability of LLM-based code completion in specific software projects based on how language, model choice, and code characteristics impact model confidence.

Authors

Institutions

Publication Details

Journal
Automated Software Engineering
Published
2026-09-30
DOI
https://doi.org/10.1007/s10515-026-00683-0
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The fools are certain; the wise are doubtful: Exploring LLM confidence in code completion

Diomidis D. Spinellis, Zoe Kotti, Konstantina Dritsa, Πάνος Λουρίδας
Automated Software Engineering
Software Engineering Research
article

The fools are certain; the wise are doubtful: Exploring LLM confidence in code completion

Diomidis D. Spinellis, Zoe Kotti, Konstantina Dritsa, Πάνος Λουρίδας
article en

Abstract

Abstract Code completion entails the task of providing missing tokens given a surrounding context. It can boost developer productivity and serve as a code discovery tool. Code completion has recently been approached with Large Language Models (LLMs) fine-tuned on code (code LLMs). The performance of code LLMs can be assessed with downstream and intrinsic metrics. Downstream metrics are usually employed to evaluate the practical utility of a model, but can be unreliable and require complex calculations and domain-specific knowledge. In contrast, intrinsic metrics such as perplexity, entropy, and mutual information, which measure model confidence or uncertainty, are simple, versatile, and universal across LLMs and tasks, and have been proposed as proxies for functional correctness and hallucination risk in LLM-generated code. Motivated by this, we evaluate the confidence of LLMs when generating code by measuring code perplexity across programming languages, models, and datasets using various LLMs, and a sample of 2254 files from 881 GitHub projects. We find that strongly-typed languages exhibit lower perplexity than dynamically typed languages. Scripting languages also demonstrate higher perplexity. Shell appears universally high in perplexity, whereas Java appears low. Code perplexity depends on the employed LLM; under a fixed model, relative language-level rankings are largely stable across evaluation corpora. Although code comments modestly increase perplexity, the language ranking based on perplexity is barely affected by their presence. LLM researchers, developers, and users can use our findings to assess the suitability of LLM-based code completion in specific software projects based on how language, model choice, and code characteristics impact model confidence.

Automated Software EngineeringVol. 33(4)
Athens University of Economics and Business (GR), Delft University of Technology (NL)
Decent work and economic growth
Openalex Percentile: Top 4%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.