Quiet or Blind? Measuring the Cost of Exclusions in LLM-Generated Detection Rules

A detection rule costs its operator twice, in analyst time for every benign alert it raises and in dwell time for every attack it misses, and the exclusion clause that encodes what is normal in one environment decides both. Language models now write such rules from threat reports, and prior work judges them by text similarity or by the attacks they catch, neither of which can see a rule that is quiet because it is blind. Here we execute the rules instead. Fifteen model configurations from seven organizations, each pinned to one endpoint and one token budget, and to one seed where the endpoint accepts it, under a frozen pre-registration, write Sigma rules for 238 SigmaHQ requirements. Every rule then runs as written, with its exclusions removed by an AST-safe transform, and beside the human rule, on 384,191 benign Windows events and 395 labeled attack files. Removing exclusions from 625 human SigmaHQ rules multiplies benign alerts 29x (196 to 5,691), and the generated rule is louder than the human on 95 of the 98 untied live comparisons on the matched requirements. Generated rules carry exclusions on 20% to 83% of the rules they produce; on the 30 requirements the benign corpus can exercise, two in five of those exclusion-bearing rules are inert, two thirds of the rest leak, and over all 1,461 exclusion-bearing rules 92% never fire. On the attack corpora the model's own exclusion is the sole reason for a miss in 93 of 1,022 exclusion-carrying cases (9.1%, 95% CI 5.8-12.9%), and 350 of 1,169 quiet generated rules (30%) are blind, covering fewer than half of the attack files the human rule catches. An arm's silence on benign data is positively rank-correlated with its blindness on attack data (Spearman rho = 0.77 over fifteen arms). Given the human's exclusion values to choose from, every configuration recovers more than it generates. A false-positive rate is therefore not interpretable without a silence rate beside it, and among the measurements reviewed here only a coverage test on attack logs separates quiet from blind. We release the harness, prompts, raw logs and pre-registrations to run it.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-13
DOI
https://doi.org/10.5281/zenodo.22733006
Primary Topic
Information and Cyber Security
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Quiet or Blind? Measuring the Cost of Exclusions in LLM-Generated Detection Rules

Umar Shoaib, Hisham Tariq
Zenodo (CERN European Organization for Nuclear Research)
Information and Cyber Security
preprint

Quiet or Blind? Measuring the Cost of Exclusions in LLM-Generated Detection Rules

Umar Shoaib, Hisham Tariq
preprint en

Abstract

A detection rule costs its operator twice, in analyst time for every benign alert it raises and in dwell time for every attack it misses, and the exclusion clause that encodes what is normal in one environment decides both. Language models now write such rules from threat reports, and prior work judges them by text similarity or by the attacks they catch, neither of which can see a rule that is quiet because it is blind. Here we execute the rules instead. Fifteen model configurations from seven organizations, each pinned to one endpoint and one token budget, and to one seed where the endpoint accepts it, under a frozen pre-registration, write Sigma rules for 238 SigmaHQ requirements. Every rule then runs as written, with its exclusions removed by an AST-safe transform, and beside the human rule, on 384,191 benign Windows events and 395 labeled attack files. Removing exclusions from 625 human SigmaHQ rules multiplies benign alerts 29x (196 to 5,691), and the generated rule is louder than the human on 95 of the 98 untied live comparisons on the matched requirements. Generated rules carry exclusions on 20% to 83% of the rules they produce; on the 30 requirements the benign corpus can exercise, two in five of those exclusion-bearing rules are inert, two thirds of the rest leak, and over all 1,461 exclusion-bearing rules 92% never fire. On the attack corpora the model's own exclusion is the sole reason for a miss in 93 of 1,022 exclusion-carrying cases (9.1%, 95% CI 5.8-12.9%), and 350 of 1,169 quiet generated rules (30%) are blind, covering fewer than half of the attack files the human rule catches. An arm's silence on benign data is positively rank-correlated with its blindness on attack data (Spearman rho = 0.77 over fifteen arms). Given the human's exclusion values to choose from, every configuration recovers more than it generates. A false-positive rate is therefore not interpretable without a silence rate beside it, and among the measurements reviewed here only a coverage test on attack logs separates quiet from blind. We release the harness, prompts, raw logs and pre-registrations to run it.

Zenodo (CERN European Organization for Nuclear Research)
Information Technology University (PK), University of Gujrat (PK)
Reduced inequalities
Information and Cyber Security
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.