Scripting WebArena: A Language-Model-Free Agent at 77.7%
LK-47.web is a WebArena agent with no language model and no learned weights. A grammar maps each task intent to one of the benchmark's 190 template identifiers and its slot values, and a hand-written plan of site skills carries the task out through the original, unmodified harness. The plans were written against the test intents, so the agent is an expert system for these templates and makes no claim to transfer to other tasks. In one run on 9 October 2026 it passed 631 of 812 tasks (77.71%), with the 84 tasks that need the harness's GPT-4 judge counted as failures. Of the 95 graded failures, 65 are tasks whose reference answer or grader disagrees with the site or the intent, and earlier work flagged 27 of them independently. Three are timing faults and 27 are the agent's own. The paper lists all 65 disputed tasks with their reasons. Code: https://github.com/LK-maker-007/LK-47.web, commit f815cd8, MIT licence. Per-task results, the runner's ledgers and all 812 trajectories: https://github.com/LK-maker-007/LK-47.web/releases/tag/run-2026-10-09 Files: the paper (PDF, 10 pages).
Authors
- Singaraj B (ORCID: https://orcid.org/0009-0002-1502-362X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-10
- DOI
- https://doi.org/10.5281/zenodo.23270749
- Primary Topic
- Web Data Mining and Analysis
- Type
- preprint