An Agency-Specific Project Authoring Advisor: LLM-Based RAG System with Automatic Prompt Optimization Method

Abstract Despite the vast potential of large language models (LLMs) in specialized fields, their application in architecture, engineering, and construction (AEC) has been hindered by several challenges, including inaccurate information retrieval, limited access to domain information, inadequate customization of prompts, and hallucination. Addressing these challenges, this paper presents an agency-specific project authoring advisor that tightly couples LLMs with a curated technical document database (TDD) in an end-to-end retrieval-augmented generation (RAG) pipeline and proposes an automated prompt optimization method within this framework. Using domain-specific embeddings, we first construct a vectorized knowledge base to enable semantic retrieval of AEC-specific information. We then employ domain-specific prompt engineering by combining persona, format template, chain of thought, and few-shot patterns to guide the model toward detailed, project authoring–specific answers. An automated adversarial prompt optimization method uses an LLM-based question generator and answer evaluation to iteratively refine prompt templates for comprehensiveness, accuracy, clarity, relevance, and conciseness. Additionally, we incorporate long-document handling using MapReduce, multiturn conversation memory, an integrated web search to cover knowledge gaps, and a chat-style web graphical user interface (GUI) for user interaction. LLM-assisted evaluation with 30 agency questions (scored for comprehensiveness, accuracy, relevance, clarity, and conciseness) showed that GPT-4 with RAG scored 75.7/100 versus 53.4/100 without RAG, and reached 88.9 with domain-specific prompt patterns and chain-of-thought reasoning. Gemini and Llama exhibited similar improvements. We also found that domain-specific embeddings deliver a 16.5% relative improvement in total averaged scores across the three LLMs. In expert validation, our advisor outperformed traditional search methods including Google and document search on all metrics, especially search efficiency. By formalizing an agency’s technical documents in a reusable knowledge base and coupling it with LLM reasoning, we contribute a transferable, end-to-end RAG pipeline for agency-specific question answering (QA), a systematic, automated prompt optimization procedure, and three reusable domain-specific prompt templates, advancing the LLM applications for knowledge-intensive and domain-specific project authoring.

Authors

Institutions

Publication Details

Journal
Journal of Computing in Civil Engineering
Published
2026-09-17
DOI
https://doi.org/10.1061/jccee5.cpeng-7192
Primary Topic
BIM and Construction Integration
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An Agency-Specific Project Authoring Advisor: LLM-Based RAG System with Automatic Prompt Optimization Method

Sebastian D. Goodfellow, Yunshun Zhong, Tamer E. El-Diraby
Journal of Computing in Civil Engineering
BIM and Construction Integration
article

An Agency-Specific Project Authoring Advisor: LLM-Based RAG System with Automatic Prompt Optimization Method

Sebastian D. Goodfellow, Yunshun Zhong, Tamer E. El-Diraby
article en

Abstract

Abstract Despite the vast potential of large language models (LLMs) in specialized fields, their application in architecture, engineering, and construction (AEC) has been hindered by several challenges, including inaccurate information retrieval, limited access to domain information, inadequate customization of prompts, and hallucination. Addressing these challenges, this paper presents an agency-specific project authoring advisor that tightly couples LLMs with a curated technical document database (TDD) in an end-to-end retrieval-augmented generation (RAG) pipeline and proposes an automated prompt optimization method within this framework. Using domain-specific embeddings, we first construct a vectorized knowledge base to enable semantic retrieval of AEC-specific information. We then employ domain-specific prompt engineering by combining persona, format template, chain of thought, and few-shot patterns to guide the model toward detailed, project authoring–specific answers. An automated adversarial prompt optimization method uses an LLM-based question generator and answer evaluation to iteratively refine prompt templates for comprehensiveness, accuracy, clarity, relevance, and conciseness. Additionally, we incorporate long-document handling using MapReduce, multiturn conversation memory, an integrated web search to cover knowledge gaps, and a chat-style web graphical user interface (GUI) for user interaction. LLM-assisted evaluation with 30 agency questions (scored for comprehensiveness, accuracy, relevance, clarity, and conciseness) showed that GPT-4 with RAG scored 75.7/100 versus 53.4/100 without RAG, and reached 88.9 with domain-specific prompt patterns and chain-of-thought reasoning. Gemini and Llama exhibited similar improvements. We also found that domain-specific embeddings deliver a 16.5% relative improvement in total averaged scores across the three LLMs. In expert validation, our advisor outperformed traditional search methods including Google and document search on all metrics, especially search efficiency. By formalizing an agency’s technical documents in a reusable knowledge base and coupling it with LLM reasoning, we contribute a transferable, end-to-end RAG pipeline for agency-specific question answering (QA), a systematic, automated prompt optimization procedure, and three reusable domain-specific prompt templates, advancing the LLM applications for knowledge-intensive and domain-specific project authoring.

Journal of Computing in Civil EngineeringVol. 41(1)
University of Toronto (CA)
Openalex Percentile: Top 14%
BIM and Construction Integration
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.