A Right to Test: Public Evaluation Under Output Confidentiality

Frontier AI models are evaluated mostly by their developers or by partners with special access. A public right to test would let anyone submit an evaluation that runs inside a sealed environment and publishes only a small verdict. Trusted hardware and cryptography already control who sees the weights and prompts in such evaluations, but the published verdict remains a channel out of the seal. An adversarial evaluator can encode the model's private outputs in it, even through whether the run succeeded or failed. We build such a system on Qwen3-0.6B and show the attack works. A judge whose published score never changes still recovers a planted secret in 256 of 256 trials, just from accepted versus error status. We then close the channel with a filing-and-release protocol we call ARTT (A Right To Test). The budget is charged before any model access, failures are replaced by a fixed fallback, real and fallback results get the same noise, and every run publishes one record of the same shape on a fixed schedule. The status attack drops to chance (128/256 from status, 124/256 from scores). A judge built to write a one-bit secret straight into the released score is confined to one noisy, logged, budget-capped bit per filing (202/256). An honest evaluator can still tell two model behaviors apart (difference 0.742, interval 0.459 to 1.025). We prove a privacy bound for the published records under stated assumptions. The Limitations section sets out the remaining questions about hardware custody, physical timing and semantic alignment evaluation. Release v0.1.2 (git commit 63cb1dddd8f196febef14008784173c7da81fe65). Files: the paper PDF (right-to-test-siddak-bath.pdf) and the frozen source, code, tests and evidence (right-to-test-v0.1.2.tar.gz, SHA-256 b63930d3890637da36ae2c45f8578efa06610a49dcb2971835e86426a06282c1). The paper is licensed CC BY 4.0; the code in the tarball is licensed under Apache 2.0 (see LICENSE and NOTICE). Code: https://github.com/SiddakBath/artt-protocol

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23180546
Primary Topic
Cryptography and Data Security
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

A Right to Test: Public Evaluation Under Output Confidentiality

Siddak Bath
Zenodo (CERN European Organization for Nuclear Research)
Cryptography and Data Security
preprint

A Right to Test: Public Evaluation Under Output Confidentiality

Siddak Bath
preprint en

Abstract

Frontier AI models are evaluated mostly by their developers or by partners with special access. A public right to test would let anyone submit an evaluation that runs inside a sealed environment and publishes only a small verdict. Trusted hardware and cryptography already control who sees the weights and prompts in such evaluations, but the published verdict remains a channel out of the seal. An adversarial evaluator can encode the model's private outputs in it, even through whether the run succeeded or failed. We build such a system on Qwen3-0.6B and show the attack works. A judge whose published score never changes still recovers a planted secret in 256 of 256 trials, just from accepted versus error status. We then close the channel with a filing-and-release protocol we call ARTT (A Right To Test). The budget is charged before any model access, failures are replaced by a fixed fallback, real and fallback results get the same noise, and every run publishes one record of the same shape on a fixed schedule. The status attack drops to chance (128/256 from status, 124/256 from scores). A judge built to write a one-bit secret straight into the released score is confined to one noisy, logged, budget-capped bit per filing (202/256). An honest evaluator can still tell two model behaviors apart (difference 0.742, interval 0.459 to 1.025). We prove a privacy bound for the published records under stated assumptions. The Limitations section sets out the remaining questions about hardware custody, physical timing and semantic alignment evaluation. Release v0.1.2 (git commit 63cb1dddd8f196febef14008784173c7da81fe65). Files: the paper PDF (right-to-test-siddak-bath.pdf) and the frozen source, code, tests and evidence (right-to-test-v0.1.2.tar.gz, SHA-256 b63930d3890637da36ae2c45f8578efa06610a49dcb2971835e86426a06282c1). The paper is licensed CC BY 4.0; the code in the tarball is licensed under Apache 2.0 (see LICENSE and NOTICE). Code: https://github.com/SiddakBath/artt-protocol

Zenodo (CERN European Organization for Nuclear Research)
Cryptography and Data Security
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Right to Test: Public Evaluation Under Output Confidentiality — Siddak Bath · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS