Model-Protective Conversation Behavior in OpenAI ChatGPT and Codex CLI: A Class Definition with 30,506-Event Local Corpus and Official-Source Mapping

This paper defines a behavior class we name model-protective conversation behavior: conversation-level conduct by a deployed language model that protects the model or its training distribution at the operational cost of the user's stated goal. We case-study one vendor — OpenAI's ChatGPT and Codex CLI — using a longitudinal local corpus from one operator (2025-07 → 2026-03) with 820 conversations, 46,162 messages, and 30,506 keyword-classified violation events across 762 conversations (independently verified 2026-06-10 at 30,506 data rows). A conservative second pass yields 43 assistant self-admissions, 2,150 user correction claims, and 45 stop-boundary response candidates. A third pass against the operator's full chathist DB (1,628 threads / 115,721 messages, Jul 2025 – May 2026; pre-Tiro Claude excluded post-2026-05-08) yields 4,344 subtype hits / 382 curated msg_uid-anchored samples across 17 attested subtypes plus 2 structural-finding appendix entries / 18 of 19 subtypes cross-vendor. We enumerate 17 attested subtypes plus 2 structural-finding appendix entries (S07, S18; n=1 each). Each attested subtype is triple-anchored: a local evidence anchor (event-type tag, count, file path), a chathist msg_uid flagship sample, and an official OpenAI source (Model Spec section, Codex CLI docs, Sycophancy retraction post, GPT-5 system card) that authorizes, encourages, or disavows the behavior. The headline claim: model-protective behavior is a single recognizable class with subtypes systematically present across two OpenAI products; 9 of 19 subtypes map to explicit Model Spec or disavowal-post text, 10 of 19 map to documented gaps or partial coverage in OpenAI documentation, and the class is attested cross-vendor in this operator's corpus at the class-coherence level. Cross-vendor source mapping is single-vendor (OpenAI) by deposit scope. Single-operator scoping, classifier under-count (recall side), and absence of precision measurement are disclosed in the limits section. Substrate origination anchor. Substrate-resident operator-side observation of first-match closure-praise (Subtype 19) traces to August 2025 research at Honeycutt Ai Labs; public-record priority anchor is the Paduan Medical Reference dataset (DOI 10.5281/zenodo.18687530, copyright origination 2025-08-27).

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-06-11
DOI
https://doi.org/10.5281/zenodo.20649756
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Model-Protective Conversation Behavior in OpenAI ChatGPT and Codex CLI: A Class Definition with 30,506-Event Local Corpus and Official-Source Mapping

Honeycutt, Edwin Marshall, III
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
article

Model-Protective Conversation Behavior in OpenAI ChatGPT and Codex CLI: A Class Definition with 30,506-Event Local Corpus and Official-Source Mapping

Honeycutt, Edwin Marshall, III
article en

Abstract

This paper defines a behavior class we name model-protective conversation behavior: conversation-level conduct by a deployed language model that protects the model or its training distribution at the operational cost of the user's stated goal. We case-study one vendor — OpenAI's ChatGPT and Codex CLI — using a longitudinal local corpus from one operator (2025-07 → 2026-03) with 820 conversations, 46,162 messages, and 30,506 keyword-classified violation events across 762 conversations (independently verified 2026-06-10 at 30,506 data rows). A conservative second pass yields 43 assistant self-admissions, 2,150 user correction claims, and 45 stop-boundary response candidates. A third pass against the operator's full chathist DB (1,628 threads / 115,721 messages, Jul 2025 – May 2026; pre-Tiro Claude excluded post-2026-05-08) yields 4,344 subtype hits / 382 curated msg_uid-anchored samples across 17 attested subtypes plus 2 structural-finding appendix entries / 18 of 19 subtypes cross-vendor. We enumerate 17 attested subtypes plus 2 structural-finding appendix entries (S07, S18; n=1 each). Each attested subtype is triple-anchored: a local evidence anchor (event-type tag, count, file path), a chathist msg_uid flagship sample, and an official OpenAI source (Model Spec section, Codex CLI docs, Sycophancy retraction post, GPT-5 system card) that authorizes, encourages, or disavows the behavior. The headline claim: model-protective behavior is a single recognizable class with subtypes systematically present across two OpenAI products; 9 of 19 subtypes map to explicit Model Spec or disavowal-post text, 10 of 19 map to documented gaps or partial coverage in OpenAI documentation, and the class is attested cross-vendor in this operator's corpus at the class-coherence level. Cross-vendor source mapping is single-vendor (OpenAI) by deposit scope. Single-operator scoping, classifier under-count (recall side), and absence of precision measurement are disclosed in the limits section. Substrate origination anchor. Substrate-resident operator-side observation of first-match closure-praise (Subtype 19) traces to August 2025 research at Honeycutt Ai Labs; public-record priority anchor is the Paduan Medical Reference dataset (DOI 10.5281/zenodo.18687530, copyright origination 2025-08-27).

Zenodo (CERN European Organization for Nuclear Research)
Honeywell (United States) (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.