Code generation for legal metadata extraction: a decomposition-based in-context learning approach

Abstract Software systems must comply with legal regulations, which is a resource-intensive task, particularly for small organizations and startups lacking dedicated legal expertise. Extracting metadata from regulations to elicit legal requirements for software is a critical step to ensure compliance. However, it is a cumbersome task due to the length and complex nature of legal text. Although prior work has pursued automated methods for extracting structural and semantic metadata from legal text, they do not consider the interplay and interrelationships among attributes associated with these metadata types, and they rely on manual labeling or heuristic-driven machine learning, which does not always generalize to new documents. In this paper, we introduce a decomposition-based in-context learning method for automatically generating a canonical representation of legal text encoded as executable Python code. Our representation is instantiated from a manually designed Python class structure that serves as a domain-specific metamodel, capturing both structural and semantic legal metadata and their interrelationships. Our corpus contains 13 US state data breach notification laws (332 paragraphs), of which six unseen laws (182 paragraphs) form the held-out test set. On this test set, our proposed method using GPT−5.1 achieves 90.5% semantic test accuracy with a precision of 79.4% and a recall of 81.9%. We also assess the generalizability of the method to the Children’s Online Privacy Protection Act (COPPA), a US federal law. The results demonstrate that, once a domain metamodel and expert-authored examples are available, few-shot code generation can extract legal metadata relationships without training a task-specific supervised model and can be adapted to unseen legislation.

Authors

Publication Details

Journal
Requirements Engineering
Published
2026-09-14
DOI
https://doi.org/10.1007/s00766-026-00468-7
Primary Topic
Artificial Intelligence in Law
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Code generation for legal metadata extraction: a decomposition-based in-context learning approach

Travis Breaux, Anmol Singhal
Requirements Engineering
Artificial Intelligence in Law
article

Code generation for legal metadata extraction: a decomposition-based in-context learning approach

Travis Breaux, Anmol Singhal
article en

Abstract

Abstract Software systems must comply with legal regulations, which is a resource-intensive task, particularly for small organizations and startups lacking dedicated legal expertise. Extracting metadata from regulations to elicit legal requirements for software is a critical step to ensure compliance. However, it is a cumbersome task due to the length and complex nature of legal text. Although prior work has pursued automated methods for extracting structural and semantic metadata from legal text, they do not consider the interplay and interrelationships among attributes associated with these metadata types, and they rely on manual labeling or heuristic-driven machine learning, which does not always generalize to new documents. In this paper, we introduce a decomposition-based in-context learning method for automatically generating a canonical representation of legal text encoded as executable Python code. Our representation is instantiated from a manually designed Python class structure that serves as a domain-specific metamodel, capturing both structural and semantic legal metadata and their interrelationships. Our corpus contains 13 US state data breach notification laws (332 paragraphs), of which six unseen laws (182 paragraphs) form the held-out test set. On this test set, our proposed method using GPT−5.1 achieves 90.5% semantic test accuracy with a precision of 79.4% and a recall of 81.9%. We also assess the generalizability of the method to the Children’s Online Privacy Protection Act (COPPA), a US federal law. The results demonstrate that, once a domain metamodel and expert-authored examples are available, few-shot code generation can extract legal metadata relationships without training a task-specific supervised model and can be adapted to unseen legislation.

Requirements EngineeringVol. 31(2)
Peace, Justice and strong institutions
Openalex Percentile: Top 3%
Artificial Intelligence in Law
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.