A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models

LLM-based systems used for tutoring can produce fluent and apparently plausible explanations, but they do not guarantee Technical Correctness, which is paramount in STEM fields such as Electrical Circuit Analysis. This study investigates how three LLM-based response systems—parametric model knowledge (Pure LLM), Retrieval-Augmented Generation (RAG) and Agentic AI—affect response quality in the field of Electrical Circuit Analysis. It also introduces the HALO Agentic AI Pipeline, a local neuro-symbolic architecture that integrates LLMs, retrieval from curated pedagogical resources and deterministic algorithms for circuit solving, and the production of verifiable artefacts (such as theoretical context, step-by-step analysis, visual representations, numerical values and pedagogical hints). We conducted a blind evaluation involving 135 students comparing Pure LLM, RAG and Agentic AI across five dimensions: Correctness, Clarity, Learning Value, Relevance and Style. The three conditions were implemented using two distinct free and open-weight LLMs: GPT-OSS:20B and Llama-3.3:70B. These were compared with ChatGPT, using GPT-5.5 Instant as a commercial benchmark. The study covered five prompts related to the Loop Current Method, seven configurations, and thirty-five anonymised responses. In addition, we conducted a technical evaluation of 72 responses generated for the 12 questions used in this study. Overall, the evaluation demonstrates that both enriched architectures (RAG and Agentic AI) were rated significantly higher than Pure LLM, with Agentic AI achieving the highest overall mean rating. Agentic AI was rated significantly higher than both alternatives in Perceived Correctness, while architecture-related differences were limited to circuit-specific prompts, where both enriched architectures outperformed Pure LLM overall, but only Agentic AI showed a significant advantage in Correctness. In the complementary technical evaluation, the Agentic AI system obtained the highest mean Technical Correctness score, compared to RAG (2.63) and Pure LLM (1.79), and only 1 of the 24 responses received a rating of 2 or lower. Additionally, for the prompts under evaluation, no statistically significant differences were detected between the highest-rated local configuration and the commercial reference, globally or across the five dimensions. Our findings shed light on the suitability of the different architectures for different tutoring tasks, from generic (merely conceptual or procedural) to circuit-specific questions. They also support the feasibility of local, tool-augmented architectures for developing tutoring applications for electrical engineering education.

Authors

Institutions

Publication Details

Journal
Computers
Published
2026-09-24
DOI
https://doi.org/10.3390/computers15100647
Primary Topic
Intelligent Tutoring Systems and Adaptive Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models

Mário Alves, João Paulo Ferreira, Armando Jorge Sousa, André Rocha et al.
Computers
Intelligent Tutoring Systems and Adaptive Learning
article

A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models

Mário Alves, João Paulo Ferreira, Armando Jorge Sousa, André Rocha, Paulo Coelho de Oliveira
article en

Abstract

LLM-based systems used for tutoring can produce fluent and apparently plausible explanations, but they do not guarantee Technical Correctness, which is paramount in STEM fields such as Electrical Circuit Analysis. This study investigates how three LLM-based response systems—parametric model knowledge (Pure LLM), Retrieval-Augmented Generation (RAG) and Agentic AI—affect response quality in the field of Electrical Circuit Analysis. It also introduces the HALO Agentic AI Pipeline, a local neuro-symbolic architecture that integrates LLMs, retrieval from curated pedagogical resources and deterministic algorithms for circuit solving, and the production of verifiable artefacts (such as theoretical context, step-by-step analysis, visual representations, numerical values and pedagogical hints). We conducted a blind evaluation involving 135 students comparing Pure LLM, RAG and Agentic AI across five dimensions: Correctness, Clarity, Learning Value, Relevance and Style. The three conditions were implemented using two distinct free and open-weight LLMs: GPT-OSS:20B and Llama-3.3:70B. These were compared with ChatGPT, using GPT-5.5 Instant as a commercial benchmark. The study covered five prompts related to the Loop Current Method, seven configurations, and thirty-five anonymised responses. In addition, we conducted a technical evaluation of 72 responses generated for the 12 questions used in this study. Overall, the evaluation demonstrates that both enriched architectures (RAG and Agentic AI) were rated significantly higher than Pure LLM, with Agentic AI achieving the highest overall mean rating. Agentic AI was rated significantly higher than both alternatives in Perceived Correctness, while architecture-related differences were limited to circuit-specific prompts, where both enriched architectures outperformed Pure LLM overall, but only Agentic AI showed a significant advantage in Correctness. In the complementary technical evaluation, the Agentic AI system obtained the highest mean Technical Correctness score, compared to RAG (2.63) and Pure LLM (1.79), and only 1 of the 24 responses received a rating of 2 or lower. Additionally, for the prompts under evaluation, no statistically significant differences were detected between the highest-rated local configuration and the commercial reference, globally or across the five dimensions. Our findings shed light on the suitability of the different architectures for different tutoring tasks, from generic (merely conceptual or procedural) to circuit-specific questions. They also support the feasibility of local, tool-augmented architectures for developing tutoring applications for electrical engineering education.

ComputersVol. 15(10)
Universidade do Porto (PT), INESC TEC (PT), Polytechnic Institute of Porto (PT)
Quality Education
Openalex Percentile: Top 9%
Intelligent Tutoring Systems and Adaptive Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.