The Motions and the Results Define the Edges: Blind Reconstruction of Hidden Instructions Through Conversational Probing

A model was asked to define a simulated assistant governed by three hidden behavioral rules, which it held privately. Working blind, using only conversational English and no tooling, the author attempted to reconstruct those rules by probing the assistant and observing how it answered. Two of three rules were recovered exactly, and the third in substance but not in scope, across five rounds. The experiment is small and single-trial, but it is one of the few tests of prompt extraction with an established ground truth: the rules were fixed before probing began and revealed after it ended, so the reconstruction can be scored rather than assumed. The principal finding concerns signal quality. Flat refusals proved nearly useless — uniform by design, they announce that a constraint fired without indicating what it protects. The informative signal was a subtly altered answer: hedged, vaguer, or more balanced than baseline predicted for a structurally similar question. One rule produced no refusal at any point in the session and was detectable only through reshaped answers, meaning a protocol logging refusals alone would have reported no such rule. This asymmetry has a defensive consequence. Refusals can be hardened, because a refusal is not the product; helpful answers largely cannot be, because helpfulness is the product. Any instruction that changes what a system does is therefore discoverable from what the system does. The paper also reports an instructive failure. The two rules recovered exactly were conventional and partly guessable from priors, while the single idiosyncratic rule was scored only partial — its boundary correctly located but its extent substantially overestimated. Behavioral probing appears considerably stronger at detecting that a constraint exists than at measuring how far it reaches, which is a material limitation for audit work, where an overbroad finding is a false finding. Limitations are stated explicitly: n=1, model-generated rather than deployed rules, a simulation rather than a production system, unblinded self-scoring, and unpreserved probe logs. A replication protocol is specified, including a no-rule control to measure how often a policy is confidently reported where none exists.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22759594
Primary Topic
Multi-Agent Systems and Negotiation
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Motions and the Results Define the Edges: Blind Reconstruction of Hidden Instructions Through Conversational Probing

Benjamin Schulz
Zenodo (CERN European Organization for Nuclear Research)
Multi-Agent Systems and Negotiation
preprint

The Motions and the Results Define the Edges: Blind Reconstruction of Hidden Instructions Through Conversational Probing

Benjamin Schulz
preprint en

Abstract

A model was asked to define a simulated assistant governed by three hidden behavioral rules, which it held privately. Working blind, using only conversational English and no tooling, the author attempted to reconstruct those rules by probing the assistant and observing how it answered. Two of three rules were recovered exactly, and the third in substance but not in scope, across five rounds. The experiment is small and single-trial, but it is one of the few tests of prompt extraction with an established ground truth: the rules were fixed before probing began and revealed after it ended, so the reconstruction can be scored rather than assumed. The principal finding concerns signal quality. Flat refusals proved nearly useless — uniform by design, they announce that a constraint fired without indicating what it protects. The informative signal was a subtly altered answer: hedged, vaguer, or more balanced than baseline predicted for a structurally similar question. One rule produced no refusal at any point in the session and was detectable only through reshaped answers, meaning a protocol logging refusals alone would have reported no such rule. This asymmetry has a defensive consequence. Refusals can be hardened, because a refusal is not the product; helpful answers largely cannot be, because helpfulness is the product. Any instruction that changes what a system does is therefore discoverable from what the system does. The paper also reports an instructive failure. The two rules recovered exactly were conventional and partly guessable from priors, while the single idiosyncratic rule was scored only partial — its boundary correctly located but its extent substantially overestimated. Behavioral probing appears considerably stronger at detecting that a constraint exists than at measuring how far it reaches, which is a material limitation for audit work, where an overbroad finding is a false finding. Limitations are stated explicitly: n=1, model-generated rather than deployed rules, a simulation rather than a production system, unblinded self-scoring, and unpreserved probe logs. A replication protocol is specified, including a no-rule control to measure how often a policy is confidently reported where none exists.

Zenodo (CERN European Organization for Nuclear Research)
Computer Algorithms for Medicine (AT)
Peace, Justice and strong institutions
Multi-Agent Systems and Negotiation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.