LLM-based Structured Intermediate Representation for building regulation encoding
Automating the encoding of building codes into formal representations is a critical step for compliance checking, yet it is often hindered by the linguistic complexity of regulations. This paper develops a two-stage pipeline that utilises Large Language Models (LLMs) to generate LegalRuleML (LRML) through a Structured Intermediate Representation (SIR). Using a dataset derived from the New Zealand Building Code (NZBC), the paper demonstrates that the SIR-based approach outperforms direct translation, with four fine-tuned models achieving F1 scores of 73.29%–77.23%, representing improvements of up to 2.40% over the current state-of-the-art method. Despite these gains, a large gap remains compared to expert-crafted SIRs (80%+). Analysis reveals that LLMs still struggle with logical structural integrity and that standard cross-entropy loss poorly aligns with formalisation goals. These findings highlight the potential of LLM-generated SIRs while underscoring the necessity of task classification and logic-driven verification in automated rule encoding systems.
Authors
- Avinash Malik (ORCID: https://orcid.org/0000-0002-7524-8292)
- Robert Amor (ORCID: https://orcid.org/0000-0002-4329-9044)
- Pinzhang Wu (ORCID: https://orcid.org/0000-0002-5511-0560)
- Partha Roop
Institutions
- University of Auckland (NZ)
Publication Details
- Journal
- Automation in Construction
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1016/j.autcon.2026.107295
- Primary Topic
- BIM and Construction Integration
- Type
- article
- Field-Weighted Citation Impact
- 0.00