A Tool-Augmented Agentic AI Pipeline for Reliable Circuit-Analysis Tutoring with Local Language Models
LLM-based systems used for tutoring can produce fluent and apparently plausible explanations, but they do not guarantee Technical Correctness, which is paramount in STEM fields such as Electrical Circuit Analysis. This study investigates how three LLM-based response systems—parametric model knowledge (Pure LLM), Retrieval-Augmented Generation (RAG) and Agentic AI—affect response quality in the field of Electrical Circuit Analysis. It also introduces the HALO Agentic AI Pipeline, a local neuro-symbolic architecture that integrates LLMs, retrieval from curated pedagogical resources and deterministic algorithms for circuit solving, and the production of verifiable artefacts (such as theoretical context, step-by-step analysis, visual representations, numerical values and pedagogical hints). We conducted a blind evaluation involving 135 students comparing Pure LLM, RAG and Agentic AI across five dimensions: Correctness, Clarity, Learning Value, Relevance and Style. The three conditions were implemented using two distinct free and open-weight LLMs: GPT-OSS:20B and Llama-3.3:70B. These were compared with ChatGPT, using GPT-5.5 Instant as a commercial benchmark. The study covered five prompts related to the Loop Current Method, seven configurations, and thirty-five anonymised responses. In addition, we conducted a technical evaluation of 72 responses generated for the 12 questions used in this study. Overall, the evaluation demonstrates that both enriched architectures (RAG and Agentic AI) were rated significantly higher than Pure LLM, with Agentic AI achieving the highest overall mean rating. Agentic AI was rated significantly higher than both alternatives in Perceived Correctness, while architecture-related differences were limited to circuit-specific prompts, where both enriched architectures outperformed Pure LLM overall, but only Agentic AI showed a significant advantage in Correctness. In the complementary technical evaluation, the Agentic AI system obtained the highest mean Technical Correctness score, compared to RAG (2.63) and Pure LLM (1.79), and only 1 of the 24 responses received a rating of 2 or lower. Additionally, for the prompts under evaluation, no statistically significant differences were detected between the highest-rated local configuration and the commercial reference, globally or across the five dimensions. Our findings shed light on the suitability of the different architectures for different tutoring tasks, from generic (merely conceptual or procedural) to circuit-specific questions. They also support the feasibility of local, tool-augmented architectures for developing tutoring applications for electrical engineering education.
Authors
- Mário Alves (ORCID: https://orcid.org/0000-0002-6139-6542)
- João Paulo Ferreira (ORCID: https://orcid.org/0000-0003-0143-9421)
- Armando Jorge Sousa (ORCID: https://orcid.org/0000-0002-0317-4714)
- André Rocha (ORCID: https://orcid.org/0000-0001-5216-7070)
- Paulo Coelho de Oliveira (ORCID: https://orcid.org/0000-0002-9982-8478)
Institutions
- Universidade do Porto (PT)
- INESC TEC (PT)
- Polytechnic Institute of Porto (PT)
Publication Details
- Journal
- Computers
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/computers15100647
- Primary Topic
- Intelligent Tutoring Systems and Adaptive Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00