Provenance by Construction: A Contract-First Architecture for Reproducible and Auditable LLM-Assisted Risk Intelligence from Heterogeneous Public Data
Analytical systems that derive risk indicators from heterogeneous public data increasingly incorporate large language models into the extraction path. This introduces a structural auditability problem: a published number may depend upon a model version, a prompt revision and a data vintage that the system never recorded, rendering the result impossible to reproduce and difficult to defend. Conventional remedies are procedural documentation conventions, review practice, post-hoc lineage reconstruction and consequently degrade under operational pressure. This paper argues for an alternative in which the required properties are enforced by the type system and validation layer, such that non-reproducible and unevidenced records are unrepresentable rather than merely discouraged. We present the architecture and contract layer of BBRI, a platform intended to derive business-risk intelligence from legally accessible Bangladeshi public data, and specify five invariants enforced mechanically at the boundary of every service: mandatory provenance with model, model-version and prompt-version identification for any language-model-derived record; append-only, vintaged observations permitting point-in-time reconstruction; a prohibition on verified claims lacking a verbatim source quotation drawn from a cited document; risk scores that carry their scoring method, configuration hash and the identifiers of every input; and source licence terms modelled as machine-readable data that gate redistribution. We describe the cross-language contract parity mechanism that keeps a hand-written TypeScript mirror synchronised with the Python definitions under continuous integration, and the legal-review protocol that precedes implementation of any data collector. The work is presented explicitly as a design contribution: the contract layer is implemented and continuously verified, while ingestion, agent orchestration, machine learning and presentation remain unimplemented, and no empirical evaluation is claimed. We set out the evaluation protocol that would substantiate the design.
Authors
- Ahnaf Akif
Institutions
- University of Dhaka (BD)
- United International University (BD)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-16
- DOI
- https://doi.org/10.5281/zenodo.22780225
- Primary Topic
- Scientific Computing and Data Management
- Type
- preprint