EvaLooop : A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming
Motivation: Evaluating the programming robustness of large language models (LLMs) is paramount for ensuring their reliability in AI-based software development. Existing robustness assessments commonly employ exogenous adversarial perturbations, where modifications are crafted externally and applied to the model input. Two limitations follow. First, different exogenous attack strategies can produce different rankings across models, making conclusions dependent on the perturbation setting. Second, these assessments do not directly characterize stability when a model's own outputs are transformed and reused as inputs. This setting is relevant to iterative software engineering workflows. Solution: We introduce EvaLooop , a novel assessment framework that evaluates robustness from a self-consistency perspective, leveraging the natural duality inherent in software engineering tasks ( e.g., code generation and code summarization). EvaLooop establishes a self-contained feedback loop where the same LLM iteratively transforms between code and natural language, producing endogenous transformations from its own output distribution. Robustness under these self-generated transformations is quantified by the Average Sustainable Loops (ASL) metric, a normalized score combining quadratically weighted successful-loop counts with semantic similarity. This cyclical strategy identifies when generated code first fails functional tests and provides a complementary view of robustness within the evaluated loops, without requiring external attack configurations. Results: We evaluate 96 popular LLMs, ranging from 0.5B to 685B parameters, on EvaLooop equipped with the MBPP Plus benchmark, and find that EvaLooop typically induces a 2.65%–47.62% absolute drop in Pass@1 accuracy within ten loops. Stability under self-generated transformations does not always align with initial performance ( i.e., a one-shot query); for instance, Qwen3-235B-A22B-Instruct-2507, despite inferior initial code generation compared to OpenAI's o-series models and DeepSeek-V3, achieves a higher ASL score. Impact: EvaLooop reveals distinct degradation patterns across different models, providing developers with a practical framework for evaluating LLM functional coherence through iterative transformations and offering complementary insights that support more informed model selection decisions. A living leaderboard: https://evalooop.github.io/ .
Authors
- Bowen Xu (ORCID: https://orcid.org/0000-0002-1006-8493)
- Sen Fang (ORCID: https://orcid.org/0000-0002-9918-7180)
- Weiyuan Ding (ORCID: https://orcid.org/0009-0006-4537-3822)
- Mengshi Zhang (ORCID: https://orcid.org/0000-0002-0025-6837)
- Zihao Chen (ORCID: https://orcid.org/0009-0004-5372-3525)
Institutions
- North Carolina State University (US)
Publication Details
- Journal
- ACM Transactions on Software Engineering and Methodology
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1145/3850240
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00