EvaLooop : A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming

Motivation: Evaluating the programming robustness of large language models (LLMs) is paramount for ensuring their reliability in AI-based software development. Existing robustness assessments commonly employ exogenous adversarial perturbations, where modifications are crafted externally and applied to the model input. Two limitations follow. First, different exogenous attack strategies can produce different rankings across models, making conclusions dependent on the perturbation setting. Second, these assessments do not directly characterize stability when a model's own outputs are transformed and reused as inputs. This setting is relevant to iterative software engineering workflows. Solution: We introduce EvaLooop , a novel assessment framework that evaluates robustness from a self-consistency perspective, leveraging the natural duality inherent in software engineering tasks ( e.g., code generation and code summarization). EvaLooop establishes a self-contained feedback loop where the same LLM iteratively transforms between code and natural language, producing endogenous transformations from its own output distribution. Robustness under these self-generated transformations is quantified by the Average Sustainable Loops (ASL) metric, a normalized score combining quadratically weighted successful-loop counts with semantic similarity. This cyclical strategy identifies when generated code first fails functional tests and provides a complementary view of robustness within the evaluated loops, without requiring external attack configurations. Results: We evaluate 96 popular LLMs, ranging from 0.5B to 685B parameters, on EvaLooop equipped with the MBPP Plus benchmark, and find that EvaLooop typically induces a 2.65%–47.62% absolute drop in Pass@1 accuracy within ten loops. Stability under self-generated transformations does not always align with initial performance ( i.e., a one-shot query); for instance, Qwen3-235B-A22B-Instruct-2507, despite inferior initial code generation compared to OpenAI's o-series models and DeepSeek-V3, achieves a higher ASL score. Impact: EvaLooop reveals distinct degradation patterns across different models, providing developers with a practical framework for evaluating LLM functional coherence through iterative transformations and offering complementary insights that support more informed model selection decisions. A living leaderboard: https://evalooop.github.io/ .

Authors

Institutions

Publication Details

Journal
ACM Transactions on Software Engineering and Methodology
Published
2026-10-08
DOI
https://doi.org/10.1145/3850240
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

EvaLooop : A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming

Bowen Xu, Sen Fang, Weiyuan Ding, Mengshi Zhang et al.
ACM Transactions on Software Engineering and Methodology
Adversarial Robustness in Machine Learning
article

EvaLooop : A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming

Bowen Xu, Sen Fang, Weiyuan Ding, Mengshi Zhang, Zihao Chen
article en

Abstract

Motivation: Evaluating the programming robustness of large language models (LLMs) is paramount for ensuring their reliability in AI-based software development. Existing robustness assessments commonly employ exogenous adversarial perturbations, where modifications are crafted externally and applied to the model input. Two limitations follow. First, different exogenous attack strategies can produce different rankings across models, making conclusions dependent on the perturbation setting. Second, these assessments do not directly characterize stability when a model's own outputs are transformed and reused as inputs. This setting is relevant to iterative software engineering workflows. Solution: We introduce EvaLooop , a novel assessment framework that evaluates robustness from a self-consistency perspective, leveraging the natural duality inherent in software engineering tasks ( e.g., code generation and code summarization). EvaLooop establishes a self-contained feedback loop where the same LLM iteratively transforms between code and natural language, producing endogenous transformations from its own output distribution. Robustness under these self-generated transformations is quantified by the Average Sustainable Loops (ASL) metric, a normalized score combining quadratically weighted successful-loop counts with semantic similarity. This cyclical strategy identifies when generated code first fails functional tests and provides a complementary view of robustness within the evaluated loops, without requiring external attack configurations. Results: We evaluate 96 popular LLMs, ranging from 0.5B to 685B parameters, on EvaLooop equipped with the MBPP Plus benchmark, and find that EvaLooop typically induces a 2.65%–47.62% absolute drop in Pass@1 accuracy within ten loops. Stability under self-generated transformations does not always align with initial performance ( i.e., a one-shot query); for instance, Qwen3-235B-A22B-Instruct-2507, despite inferior initial code generation compared to OpenAI's o-series models and DeepSeek-V3, achieves a higher ASL score. Impact: EvaLooop reveals distinct degradation patterns across different models, providing developers with a practical framework for evaluating LLM functional coherence through iterative transformations and offering complementary insights that support more informed model selection decisions. A living leaderboard: https://evalooop.github.io/ .

ACM Transactions on Software Engineering and Methodology
North Carolina State University (US)
Openalex Percentile: Top 12%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.