Supporting Multi-Step Workflows with LLM-Based Web Automation using Agentic Context Engineering and Selective Multimodal Reasoning

[Background.] Multi-step web workflows, such as comparing offers across portals, cost people much time. Scripts must be written for each site and break when pages change; agents built on large language models can pursue a goal stated in natural language on unseen sites. Current agents depend on site-specific interaction, spend many tokens on large pages and redundant steps, use expensive vision models at every step, do not reuse earlier work, and give users few safeguards. [Objective.] This thesis develops and evaluates a generalized, trustworthy and human-centred web agent for enterprise-relevant tasks, with five objectives: generalization, a bounded cost, intent reasoning and task decomposition, knowing when to involve a person, and safety, security and privacy. [Approach.] Each limitation leads to a design principle and a component. Structured DOM preprocessing represents any page as site-independent records, notes what it could not read in a coverage manifest, and gives the model only a compact extract. URL optimization shortens interactions, and parallel agents handle the subtasks of a decomposed goal. Selective multimodal reasoning calls a vision model only when text is not enough. Agentic context engineering reuses experience through curated playbooks and a memory of earlier runs. A policy layer separates permitting an action from executing it and hands control to a person at logins and approvals, and a completion is accepted only when stored evidence supports it. [Evaluation.] On 1,656 scored tasks from WebVoyager, Online-Mind2Web, WebArena and Odysseys, run with grok-4-1-fast-reasoning (Grok) on live sites and DeepSeek-V4-Flash on WebArena, strict success is 69.7% to 95.9%. On the 229 Online-Mind2Web tasks every variant reached, the full architecture on Grok solves 86.5% against 65.1% for the same model given only the DOM and a screenshot, and 69.4% against 40.8% of the hard ones, at about six times the cost per completed task. Without its completion check, wrong completions rise from 2.8% to 15.3%; without vision, success falls by 11 points. Under Grok a task costs $0.10–$0.20 for routine browsing and up to $1.54 for long research ($3.56 under DeepSeek); GPT-6.1 Sol matches Grok at six times the cost. In a pilot, six people checked a saved run in a median of 3.5 minutes and accepted 10 of 12 answers unchanged. [Conclusion.] A suitable architecture lets an inexpensive model succeed on multistep web tasks across very different sites, where the same model without it solves far fewer. Evidence handling, safety policy and human control belong to the runtime, so they carry over to stronger models. Left open are the effect of memory alone, false completions on long tasks, and attacks on the safeguards.

Authors

Publication Details

Journal
Leibniz Universität Hannover
Published
2026-10-08
DOI
https://doi.org/10.15488/22777
Primary Topic
Web Data Mining and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Supporting Multi-Step Workflows with LLM-Based Web Automation using Agentic Context Engineering and Selective Multimodal Reasoning

Amirreza Alasti
Leibniz Universität Hannover
Web Data Mining and Analysis
article

Supporting Multi-Step Workflows with LLM-Based Web Automation using Agentic Context Engineering and Selective Multimodal Reasoning

Amirreza Alasti
article en

Abstract

[Background.] Multi-step web workflows, such as comparing offers across portals, cost people much time. Scripts must be written for each site and break when pages change; agents built on large language models can pursue a goal stated in natural language on unseen sites. Current agents depend on site-specific interaction, spend many tokens on large pages and redundant steps, use expensive vision models at every step, do not reuse earlier work, and give users few safeguards. [Objective.] This thesis develops and evaluates a generalized, trustworthy and human-centred web agent for enterprise-relevant tasks, with five objectives: generalization, a bounded cost, intent reasoning and task decomposition, knowing when to involve a person, and safety, security and privacy. [Approach.] Each limitation leads to a design principle and a component. Structured DOM preprocessing represents any page as site-independent records, notes what it could not read in a coverage manifest, and gives the model only a compact extract. URL optimization shortens interactions, and parallel agents handle the subtasks of a decomposed goal. Selective multimodal reasoning calls a vision model only when text is not enough. Agentic context engineering reuses experience through curated playbooks and a memory of earlier runs. A policy layer separates permitting an action from executing it and hands control to a person at logins and approvals, and a completion is accepted only when stored evidence supports it. [Evaluation.] On 1,656 scored tasks from WebVoyager, Online-Mind2Web, WebArena and Odysseys, run with grok-4-1-fast-reasoning (Grok) on live sites and DeepSeek-V4-Flash on WebArena, strict success is 69.7% to 95.9%. On the 229 Online-Mind2Web tasks every variant reached, the full architecture on Grok solves 86.5% against 65.1% for the same model given only the DOM and a screenshot, and 69.4% against 40.8% of the hard ones, at about six times the cost per completed task. Without its completion check, wrong completions rise from 2.8% to 15.3%; without vision, success falls by 11 points. Under Grok a task costs $0.10–$0.20 for routine browsing and up to $1.54 for long research ($3.56 under DeepSeek); GPT-6.1 Sol matches Grok at six times the cost. In a pilot, six people checked a saved run in a median of 3.5 minutes and accepted 10 of 12 answers unchanged. [Conclusion.] A suitable architecture lets an inexpensive model succeed on multistep web tasks across very different sites, where the same model without it solves far fewer. Evidence handling, safety policy and human control belong to the runtime, so they carry over to stronger models. Left open are the effect of memory alone, false completions on long tasks, and attacks on the safeguards.

Leibniz Universität Hannover
Openalex Percentile: Top 6%
Web Data Mining and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.