LLMutantKiller: Using Large Language Models to Generate Tests That Kill Mutants

The primary goal of mutation testing is to assess the quality of an application’s test suite. This is accomplished by introducing syntactic changes into a program and determining if any test failures occur for the resulting mutated program, commonly referred to as a mutant. If so, the mutant is said to be killed, confirming that the test suite is of sufficient quality to detect the introduced fault. A problem arises if a mutant does not impact the behavior of any test. Such a surviving mutant may occur for two reasons: either it involves a semantics- preserving program transformation or the test suite is not strong enough. Determining why a mutant survives often involves complex, non-local reasoning. This paper presents an LLM-based test generation technique for killing surviving mutants, implemented in a tool called LLMutantKiller. The technique is feedback-directed in the sense that if a test is produced that does not kill a given mutant, the LLM is re-prompted up to a specified number of times with scenario-specific feedback such as syntax errors, dependency violations, or execution logs (e.g., failing assertions) and asked to try again. We evaluate LLMutantKiller on 915 randomly selected surviving mutants produced by StrykerJS, a state-of-the-art mutation testing tool, across 13 open-source JavaScript/TypeScript applications. The results show that LLMutantKiller kills up to 95.3% of the surviving mutants classified as inducing behavioral changes and that it rarely produces invalid tests.

Authors

Institutions

Publication Details

Journal
Proceedings of the ACM on software engineering.
Published
2026-10-01
DOI
https://doi.org/10.1145/3832098
Primary Topic
Software Testing and Debugging Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

LLMutantKiller: Using Large Language Models to Generate Tests That Kill Mutants

Farideh Khalili, Frank Tip, Harshit Garg, Aidan Domondon
Proceedings of the ACM on software engineering.
Software Testing and Debugging Techniques
article

LLMutantKiller: Using Large Language Models to Generate Tests That Kill Mutants

Farideh Khalili, Frank Tip, Harshit Garg, Aidan Domondon
article en

Abstract

The primary goal of mutation testing is to assess the quality of an application’s test suite. This is accomplished by introducing syntactic changes into a program and determining if any test failures occur for the resulting mutated program, commonly referred to as a mutant. If so, the mutant is said to be killed, confirming that the test suite is of sufficient quality to detect the introduced fault. A problem arises if a mutant does not impact the behavior of any test. Such a surviving mutant may occur for two reasons: either it involves a semantics- preserving program transformation or the test suite is not strong enough. Determining why a mutant survives often involves complex, non-local reasoning. This paper presents an LLM-based test generation technique for killing surviving mutants, implemented in a tool called LLMutantKiller. The technique is feedback-directed in the sense that if a test is produced that does not kill a given mutant, the LLM is re-prompted up to a specified number of times with scenario-specific feedback such as syntax errors, dependency violations, or execution logs (e.g., failing assertions) and asked to try again. We evaluate LLMutantKiller on 915 randomly selected surviving mutants produced by StrykerJS, a state-of-the-art mutation testing tool, across 13 open-source JavaScript/TypeScript applications. The results show that LLMutantKiller kills up to 95.3% of the surviving mutants classified as inducing behavioral changes and that it rarely produces invalid tests.

Proceedings of the ACM on software engineering.Vol. 3(ISSTA)
Northeastern University (US), Amazon (United States) (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 7%
Software Testing and Debugging Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

LLMutantKiller: Using Large Language Models to Generate Tests That Kill Mutants — Farideh Khalili, Frank Tip, et al. · Proceedings of the ACM on software engineering. (2026) | TGRS Research Map | TGRS