Evaluating Methodological Skills in Public Statistics: an exploratory study of the contribution of institutional knowledge to the behavior of language-model agents
Incorporating procedural knowledge through skills expands the possibilities for specializing agents built on language models. This article asks whether the methodological skills built by a public statistics institution from its own published body of work improve the responses of the agent that loads them and, if so, through what mechanism. We present a retrospective study of 16 series of evaluations conducted within the project to build Fundação Seade's agentic and institutional infrastructure, using agent profiles cloned from production and experimental conditions that differ by a single variable, spanning fertility, nuptiality, the labor market, regional development, and foreign trade. Five consolidated series comprise 97 recorded runs, and additional series add 56 documented outputs. Three results organize the analysis. First, the skill's gain is localized and methodological in nature: on direct factual questions the conditions tie, and the difference emerges when the question contains a methodological trap, such as deriving a territorial breakdown by subtraction or confusing age composition with reproductive timing. Second, the skill's content acts unevenly: institutional convention is transmitted (3/3 versus 0/3), a prohibition offered without an alternative leaks through (1/4 before and 0/12 after naming the legitimate outputs), and dated values are not reproduced when the source is accessible (0 of 10 exclusive values in 12 responses). Third, the presence of the skill does not protect against extrapolation: in foreign trade, responses with the skill inferred volume from weight more often than responses without it. We propose classifying skill content by function — reference-frame structure, decision-changing parameter, and dated value — and a behavioral evaluation protocol based on questions with methodological traps, judgment under hidden labels, and tool-call tracing. The article closes with recommendations for building skills in public-statistics products. The heterogeneity of the designs, the sample sizes, and the partially blinded nature of the evaluations restrict generalization.
Authors
- Vagner Bessa
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23182680
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint