Evaluating Methodological Skills in Public Statistics: an exploratory study of the contribution of institutional knowledge to the behavior of language-model agents

Incorporating procedural knowledge through skills expands the possibilities for specializing agents built on language models. This article asks whether the methodological skills built by a public statistics institution from its own published body of work improve the responses of the agent that loads them and, if so, through what mechanism. We present a retrospective study of 16 series of evaluations conducted within the project to build Fundação Seade's agentic and institutional infrastructure, using agent profiles cloned from production and experimental conditions that differ by a single variable, spanning fertility, nuptiality, the labor market, regional development, and foreign trade. Five consolidated series comprise 97 recorded runs, and additional series add 56 documented outputs. Three results organize the analysis. First, the skill's gain is localized and methodological in nature: on direct factual questions the conditions tie, and the difference emerges when the question contains a methodological trap, such as deriving a territorial breakdown by subtraction or confusing age composition with reproductive timing. Second, the skill's content acts unevenly: institutional convention is transmitted (3/3 versus 0/3), a prohibition offered without an alternative leaks through (1/4 before and 0/12 after naming the legitimate outputs), and dated values are not reproduced when the source is accessible (0 of 10 exclusive values in 12 responses). Third, the presence of the skill does not protect against extrapolation: in foreign trade, responses with the skill inferred volume from weight more often than responses without it. We propose classifying skill content by function — reference-frame structure, decision-changing parameter, and dated value — and a behavioral evaluation protocol based on questions with methodological traps, judgment under hidden labels, and tool-call tracing. The article closes with recommendations for building skills in public-statistics products. The heterogeneity of the designs, the sample sizes, and the partially blinded nature of the evaluations restrict generalization.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23182680
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Evaluating Methodological Skills in Public Statistics: an exploratory study of the contribution of institutional knowledge to the behavior of language-model agents

Vagner Bessa
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Evaluating Methodological Skills in Public Statistics: an exploratory study of the contribution of institutional knowledge to the behavior of language-model agents

Vagner Bessa
preprint en

Abstract

Incorporating procedural knowledge through skills expands the possibilities for specializing agents built on language models. This article asks whether the methodological skills built by a public statistics institution from its own published body of work improve the responses of the agent that loads them and, if so, through what mechanism. We present a retrospective study of 16 series of evaluations conducted within the project to build Fundação Seade's agentic and institutional infrastructure, using agent profiles cloned from production and experimental conditions that differ by a single variable, spanning fertility, nuptiality, the labor market, regional development, and foreign trade. Five consolidated series comprise 97 recorded runs, and additional series add 56 documented outputs. Three results organize the analysis. First, the skill's gain is localized and methodological in nature: on direct factual questions the conditions tie, and the difference emerges when the question contains a methodological trap, such as deriving a territorial breakdown by subtraction or confusing age composition with reproductive timing. Second, the skill's content acts unevenly: institutional convention is transmitted (3/3 versus 0/3), a prohibition offered without an alternative leaks through (1/4 before and 0/12 after naming the legitimate outputs), and dated values are not reproduced when the source is accessible (0 of 10 exclusive values in 12 responses). Third, the presence of the skill does not protect against extrapolation: in foreign trade, responses with the skill inferred volume from weight more often than responses without it. We propose classifying skill content by function — reference-frame structure, decision-changing parameter, and dated value — and a behavioral evaluation protocol based on questions with methodological traps, judgment under hidden labels, and tool-call tracing. The article closes with recommendations for building skills in public-statistics products. The heterogeneity of the designs, the sample sizes, and the partially blinded nature of the evaluations restrict generalization.

Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.