MAESTRO: Fine-Tuned Multiagent Large Action Model-Based GenAI Assistant Integrating IoT Sensing, Smart Control, and Extended Reality for Human-Building Interaction
Abstract There has been recent interest in smart human–building interaction (HBI) interfaces to improve how occupants interact with building systems. Intelligent facility management (FM) utilizes HBI principles to transform static infrastructures into dynamic, human-centric environments, where facility managers need to interpret live Internet-of-things (IoT) sensor data, consult procedures, and execute safe control actions for various building systems. These tasks are fragmented across dashboards, manuals, and control panels, increasing cognitive load and slowing down decision-making. To address these challenges, this study proposes the multiagent extended-reality and sensing-aware task-level reasoning orchestrator (MAESTRO). MAESTRO is a multimodal agentic GenAI assistant comprised of four agents that support procedural queries, telemetry interpretation, and supervised device actuation in a single HBI loop. MAESTRO integrates a large action model (LAM)-based system with retrieval-augmented reasoning interfaced in an extended reality (XR) environment and mediated by a Model Context Protocol. It uses an orchestrator, which is the coordination component that interprets each user request and decides which agent capability should handle the request, while enforcing supervised and schema validated execution for sensing and actuation. The orchestrator routes each request to the appropriate agents including a domain-adapted FM language agent trained on manuals, logs, and work orders; an IoT control agent that generates schema validated tool calls to sensors and actuators; and a vision language agent for image-based queries. MAESTRO was evaluated on various HBI testing scenarios covering safety inspection, system control, troubleshooting, and procedural guidance. The results show improved language modeling and answer quality after adaptation, a valid tool-call rate of 91%, device-execution success of 89%, telemetry match rate of 81%, and an overall end-to-end task success rate exceeding 84%. This paper shows the potential of integrating fine-tuned LAM-based agents with structured IoT-enabled middleware and XR-based HBI systems, demonstrating how foundation-model reasoning and actuation can support more intelligent and adaptive HBI workflows.
Authors
- Rayan H. Assaad (ORCID: https://orcid.org/0000-0003-4626-5656)
- Andrew B. Park (ORCID: https://orcid.org/0009-0005-4192-4159)
- Oscar Poudel (ORCID: https://orcid.org/0009-0007-0643-7974)
- Mohamad Awada
Institutions
- New Jersey Institute of Technology (US)
- New York University (US)
Publication Details
- Journal
- Journal of Computing in Civil Engineering
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1061/jccee5.cpeng-7915
- Primary Topic
- BIM and Construction Integration
- Type
- article
- Field-Weighted Citation Impact
- 0.00