Withdrawing and Refusing Post-Deployment Learning Claims: An Instrumented Case Study of Prólogo, a Persistent LLM Assistant System

Claims that a deployed language-model agent has learned can fail in more than one way. An experience artifact may exist without being delivered to the model, a desired behavior may occur without a discriminating counterfactual, and a runtime may enforce a property that the model never learned. We report an instrumented longitudinal case study of Prólogo, a persistent LLM assistant system whose post-deployment learning claims were accepted, withdrawn, or refused as these evidence dimensions disagreed. We retain the historical withdrawal of negative efficacy interpretations; its original delivery finding is not independently established by the current reconstruction. Historical evaluations reported delivery as passed in two single-consumption same-task executions that also showed the pre-specified recovery behavior; neither observation provided causal attribution. A subsequent frozen no-lesson control and lesson-bearing treatment cohort was operationally valid yet non-discriminating: both arms recovered in 3/3 runs and the cohort was classified as inadmissible. Three sequential, adaptively designed control-only calibration episodes exposed distinct instrument defects. The terminal V3 episode had four structurally admissible runs, objective recovery in 1/4, bounded completion in 0/4, and recovery by sweep in 1/4; its frozen classification was cohort_unsuitable_sweep_dominant. We report no causal learning or transfer result. The contribution is a case-derived claim-analysis schema that separates artifact existence, release attribution, consumed-context delivery, behavior, counterfactual discriminability, recurrence, transfer, persistence, and deterministic enforcement. In the observed lineage, an evidence and adjudication structure and an operator-mediated governed process enabled earlier interpretations to be withdrawn or refused without rewriting their historical traces. External human evidence review remains outstanding.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22882050
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Withdrawing and Refusing Post-Deployment Learning Claims: An Instrumented Case Study of Prólogo, a Persistent LLM Assistant System

Alexandre Lima Gomes
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

Withdrawing and Refusing Post-Deployment Learning Claims: An Instrumented Case Study of Prólogo, a Persistent LLM Assistant System

Alexandre Lima Gomes
preprint en

Abstract

Claims that a deployed language-model agent has learned can fail in more than one way. An experience artifact may exist without being delivered to the model, a desired behavior may occur without a discriminating counterfactual, and a runtime may enforce a property that the model never learned. We report an instrumented longitudinal case study of Prólogo, a persistent LLM assistant system whose post-deployment learning claims were accepted, withdrawn, or refused as these evidence dimensions disagreed. We retain the historical withdrawal of negative efficacy interpretations; its original delivery finding is not independently established by the current reconstruction. Historical evaluations reported delivery as passed in two single-consumption same-task executions that also showed the pre-specified recovery behavior; neither observation provided causal attribution. A subsequent frozen no-lesson control and lesson-bearing treatment cohort was operationally valid yet non-discriminating: both arms recovered in 3/3 runs and the cohort was classified as inadmissible. Three sequential, adaptively designed control-only calibration episodes exposed distinct instrument defects. The terminal V3 episode had four structurally admissible runs, objective recovery in 1/4, bounded completion in 0/4, and recovery by sweep in 1/4; its frozen classification was cohort_unsuitable_sweep_dominant. We report no causal learning or transfer result. The contribution is a case-derived claim-analysis schema that separates artifact existence, release attribution, consumed-context delivery, behavior, counterfactual discriminability, recurrence, transfer, persistence, and deterministic enforcement. In the observed lineage, an evidence and adjudication structure and an operator-mediated governed process enabled earlier interpretations to be withdrawn or refused without rewriting their historical traces. External human evidence review remains outstanding.

Zenodo (CERN European Organization for Nuclear Research)
Reduced inequalities
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.