When 99% Means Nothing: Negative Results from Measuring Semantic Equivalence Between Agent Instructions
Multi-agent LLM systems pass instructions between agents, and a rewrite that silently changes meaning ("delete" for "archive", a dropped negation) can trigger irreversible actions. We built AIXL, a compact semantic interchange format with a closed vocabulary (25 actions, about 20 nouns; Spanish, English, Portuguese), a 16-dimension canonical form, and a deterministic comparator, hoping it would let agents verify that two instructions mean the same. On our own 500-pair benchmark it scored 99.4%. On blind sets written by other language models without access to the code, the rule-based route scored 25.3% and 37.3% on the two independent sets, and about 60% of non-equivalent pairs were judged equivalent, because the format silently discards what it has no slot for. Five successive lines of repair (format-level safeguards, LLM-assisted routes, a closed-domain semantic track, an atom layer, and a distilled local translator) reduced false equivalences but none could prove equivalence on open text; three automated labelers produced two or more equivalent encodings for only 5.6% of 700 freshly written instructions, which caps any label-trained extractor. What survived is modest: a full-text arbiter in which two independent LLM judges must both say "same" (100% of equivalents proven, 0.6% false on one set; 0 of 480 false on near-miss negatives on another) is the measured baseline, and AIXL is useful as a loss detector, not as a prover. We report every result, including pre-registered gates that failed, and distill methodological lessons on self-authored benchmarks, spent test sets, and label stability.
Authors
- Darwing John Pérez Aranguren (ORCID: https://orcid.org/0009-0007-5187-5295)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23227168
- Primary Topic
- Natural Language Processing Techniques
- Type
- preprint