Scripting WebArena: A Language-Model-Free Agent at 77.7%

LK-47.web is a WebArena agent with no language model and no learned weights. A grammar maps each task intent to one of the benchmark's 190 template identifiers and its slot values, and a hand-written plan of site skills carries the task out through the original, unmodified harness. The plans were written against the test intents, so the agent is an expert system for these templates and makes no claim to transfer to other tasks. In one run on 9 October 2026 it passed 631 of 812 tasks (77.71%), with the 84 tasks that need the harness's GPT-4 judge counted as failures. Of the 95 graded failures, 65 are tasks whose reference answer or grader disagrees with the site or the intent, and earlier work flagged 27 of them independently. Three are timing faults and 27 are the agent's own. The paper lists all 65 disputed tasks with their reasons. Code: https://github.com/LK-maker-007/LK-47.web, commit f815cd8, MIT licence. Per-task results, the runner's ledgers and all 812 trajectories: https://github.com/LK-maker-007/LK-47.web/releases/tag/run-2026-10-09 Files: the paper (PDF, 10 pages).

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-10
DOI
https://doi.org/10.5281/zenodo.23270749
Primary Topic
Web Data Mining and Analysis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Scripting WebArena: A Language-Model-Free Agent at 77.7%

Singaraj B
Zenodo (CERN European Organization for Nuclear Research)
Web Data Mining and Analysis
preprint

Scripting WebArena: A Language-Model-Free Agent at 77.7%

Singaraj B
preprint en

Abstract

LK-47.web is a WebArena agent with no language model and no learned weights. A grammar maps each task intent to one of the benchmark's 190 template identifiers and its slot values, and a hand-written plan of site skills carries the task out through the original, unmodified harness. The plans were written against the test intents, so the agent is an expert system for these templates and makes no claim to transfer to other tasks. In one run on 9 October 2026 it passed 631 of 812 tasks (77.71%), with the 84 tasks that need the harness's GPT-4 judge counted as failures. Of the 95 graded failures, 65 are tasks whose reference answer or grader disagrees with the site or the intent, and earlier work flagged 27 of them independently. Three are timing faults and 27 are the agent's own. The paper lists all 65 disputed tasks with their reasons. Code: https://github.com/LK-maker-007/LK-47.web, commit f815cd8, MIT licence. Per-task results, the runner's ledgers and all 812 trajectories: https://github.com/LK-maker-007/LK-47.web/releases/tag/run-2026-10-09 Files: the paper (PDF, 10 pages).

Zenodo (CERN European Organization for Nuclear Research)
Web Data Mining and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Scripting WebArena: A Language-Model-Free Agent at 77.7% — Singaraj B · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS