The Agentic Brain: Dual-Process Metacognition, Evidence-Gated Authority, and Runtime Governance
As AI agents move from producing answers to taking consequential actions, they must decide both how much reasoning to perform and what evidence justifies acting. We develop the agentic control plane, joining a System-1-like fast proposer, a System-2-like portfolio of deliberation and verification, a metacognitive router, and an effect-aware commitment gate. The central coupling is that evidence changes not only beliefs but also the actions the agent is permitted to take. The agent therefore optimizes over part of the process that governs its own authority. We formalize this interaction as resource-bounded metareasoning in a shielded partially observable decision process. A four-corner value-of-computation representation distinguishes belief changes, admissibility changes, and their interaction. It exposes complementary evidence chains that defeat myopic routing and identifies conditions under which governance makes useful information privately costly. Controlled counterfactuals distinguish strategic ignorance—avoiding a check because its result could remove permission—from ordinary cost-sensitive skipping. This analysis explains how the same deliberative capability can support verification or exploit weak evidence contracts. From verifier bounds covering the permitted adaptive selection process, we construct statistical evidence and bound unsafe admission across adaptive testing, stopping, and proposal generation. A composition theorem separates shared calibration uncertainty, proposal-level risk spending, and risk from selected fast-path uses. A reference authorization mechanism binds evidence to the executed payload, target and policy versions, dependencies, and expiry, with protected workflow identities, atomic risk reservations, and single-use authorizations. The learning analysis further separates competence from permission: distilling costly deliberation into fast behavior does not automatically transfer the evidence required for authorized use. The theory is instantiated in an exactly solved finite environment with heterogeneous verification, candidate revision, resource constraints, trained tabular routers, and checked execution. In a bounded adversarial test of a selectable verifier blind spot, a deliberately misspecified pooled evidence model remains fully exploitable despite complete logging; selection-valid evidence reduces maximum unsafe admission from 100% to 5%, and a required scan reduces it to 1.8%. In one aligned-objective comparison, forecasting future permission raises synthetic principal utility from 4.33 to 6.04 and lowers runtime cost, with a small increase in realized unsafe commitment. When objectives differ, foresight can instead lower principal value. Governance-response interventions expose both evidence avoidance and null or reversed effects. A planning-depth sweep separates insufficient lookahead from learning advantage, while exhaustive finite-sample calibration quantifies the performance cost of estimated verifier bounds. Further verification can remain optimal even after permission is available. Together, the formal results and executable construction establish an incentive-and-assurance framework for evidence-gated agency within a declared finite model. The broader implication is that agentic safety requires objective alignment, reliable evidence acquisition, and protected control of external effects to operate together. Dual-process cognition explains why reasoning must be allocated; evidence-gated control explains how that allocation becomes consequential for alignment and safety. Keywords: agentic AI; dual-process cognition; System 1; System 2; metacognition; metareasoning; evidence-gated authority; runtime governance; value of computation; verification; information flow; reasoning distillation.
Authors
- Fitih M. Cinnor
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23125975
- Primary Topic
- Scientific Computing and Data Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00