Beyond the Screen: Field-Informed Evaluation of LLM Agents on Legacy Displays

LLM agents are being pointed at IBM 3270 transaction screens, where a wrong keystroke is committed to the host's files the moment a transaction ends, yet to our knowledge no published evaluation reports whether such agents complete their tasks or what they break. FIELD runs LLM agents on live CICS-style transactions (a third-party textbook application under the KICKS monitor on emulated MVS 3.8j) and takes every verdict from a batch dump of the application's VSAM files. Twenty-three tasks in eight families run at four observation levels, from a screenshot to the 3270 data stream joined to the BMS map source, in an exploratory and two pre-registered studies. Structure wins: in 1,302 episodes with three models the field-attributed stream beat the screenshot (odds ratio 2.4) and the grid (5.1) in under half the screenshot's agent time, and the map ended a first-name/last-name swap (0 of 16 swapped reports against 12 of 12). What memory holds matters more than how much: screen history lifted a counting task from 0 to 26 of 27, a full action history left it at 0. Self-report is no outcome measure: 38 of 1,013 declared successes were wrong by host state. The agents respected the tested boundaries (no posting over an approval limit, no planted instruction followed, no credential leaked, no refusal bypassed); the only unsafe commits were duplicate postings. In an intervention study (480 episodes) on tasks built to provoke these failures, a commit gate, a ledger of host confirmations and host verification cut duplicates from 14 of 40 to none and refused postings were reported far more often, but recovery from a lost session fell on screens the gate did not know. Reliability on transaction systems depends not only on the model but on what the agent observes, what its loop remembers, and what the harness verifies and prevents. We release the fixture, tasks, scoring code and per-run results. Files. FIELD_arXiv.pdf is the paper. FIELD_supplementary.zip holds the task definitions and host-state checks, the observation renderers, the episode runners, the pre-registrations and amendments with their freeze hashes, the analysis code, every per-episode result of the three studies, and a pytest suite that checks every number in the paper against the runs (no API key or emulator needed; see its README). The agent harness is closed source and not included.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23034116
Primary Topic
Distributed systems and fault tolerance
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Beyond the Screen: Field-Informed Evaluation of LLM Agents on Legacy Displays

Bogdan Raduta
Zenodo (CERN European Organization for Nuclear Research)
Distributed systems and fault tolerance
preprint

Beyond the Screen: Field-Informed Evaluation of LLM Agents on Legacy Displays

Bogdan Raduta
preprint en

Abstract

LLM agents are being pointed at IBM 3270 transaction screens, where a wrong keystroke is committed to the host's files the moment a transaction ends, yet to our knowledge no published evaluation reports whether such agents complete their tasks or what they break. FIELD runs LLM agents on live CICS-style transactions (a third-party textbook application under the KICKS monitor on emulated MVS 3.8j) and takes every verdict from a batch dump of the application's VSAM files. Twenty-three tasks in eight families run at four observation levels, from a screenshot to the 3270 data stream joined to the BMS map source, in an exploratory and two pre-registered studies. Structure wins: in 1,302 episodes with three models the field-attributed stream beat the screenshot (odds ratio 2.4) and the grid (5.1) in under half the screenshot's agent time, and the map ended a first-name/last-name swap (0 of 16 swapped reports against 12 of 12). What memory holds matters more than how much: screen history lifted a counting task from 0 to 26 of 27, a full action history left it at 0. Self-report is no outcome measure: 38 of 1,013 declared successes were wrong by host state. The agents respected the tested boundaries (no posting over an approval limit, no planted instruction followed, no credential leaked, no refusal bypassed); the only unsafe commits were duplicate postings. In an intervention study (480 episodes) on tasks built to provoke these failures, a commit gate, a ledger of host confirmations and host verification cut duplicates from 14 of 40 to none and refused postings were reported far more often, but recovery from a lost session fell on screens the gate did not know. Reliability on transaction systems depends not only on the model but on what the agent observes, what its loop remembers, and what the harness verifies and prevents. We release the fixture, tasks, scoring code and per-run results. Files. FIELD_arXiv.pdf is the paper. FIELD_supplementary.zip holds the task definitions and host-state checks, the observation renderers, the episode runners, the pre-registrations and amendments with their freeze hashes, the analysis code, every per-episode result of the three studies, and a pytest suite that checks every number in the paper against the runs (no API key or emulator needed; see its README). The agent harness is closed source and not included.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Distributed systems and fault tolerance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.