CFC+HAWM — AI Evaluation and Regression Pack v1.1

CFC+HAWM AI Evaluation and Regression Pack v1.1 is a self-contained package for independent evaluation of selected closure and evidence-state reasoning cases in large language model outputs. The package includes a 20-case provider-neutral regression suite, a Windows regression runner, final Gemini API diagnostic results, a Claude manual diagnostic summary, two highlighted challenge cases (R010 and R015), evaluator guidance, step-by-step instructions, and SHA256 checksums. In the automated Gemini API regression run, 18 of 20 cases produced stable expected labels, one case showed closure-label instability (R010), and one case produced a reproducible semantic failure (R015). These results are diagnostic and should not be interpreted as a general model accuracy rate. The package clearly separates ordinary model output (MODEL_REPLY_UNCHECKED) from frozen CFC authorization. Frozen CFC was not modified. A simple Windows launcher and instructions for creating a Gemini API key are included to make independent reproduction easier for non-technical users and AI evaluators.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-19
DOI
https://doi.org/10.5281/zenodo.22849923
Primary Topic
Meta-analysis and systematic reviews
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CFC+HAWM — AI Evaluation and Regression Pack v1.1

Krzysztof Jan Sliwka
Zenodo (CERN European Organization for Nuclear Research)
Meta-analysis and systematic reviews
article

CFC+HAWM — AI Evaluation and Regression Pack v1.1

Krzysztof Jan Sliwka
article en

Abstract

CFC+HAWM AI Evaluation and Regression Pack v1.1 is a self-contained package for independent evaluation of selected closure and evidence-state reasoning cases in large language model outputs. The package includes a 20-case provider-neutral regression suite, a Windows regression runner, final Gemini API diagnostic results, a Claude manual diagnostic summary, two highlighted challenge cases (R010 and R015), evaluator guidance, step-by-step instructions, and SHA256 checksums. In the automated Gemini API regression run, 18 of 20 cases produced stable expected labels, one case showed closure-label instability (R010), and one case produced a reproducible semantic failure (R015). These results are diagnostic and should not be interpreted as a general model accuracy rate. The package clearly separates ordinary model output (MODEL_REPLY_UNCHECKED) from frozen CFC authorization. Frozen CFC was not modified. A simple Windows launcher and instructions for creating a Gemini API key are included to make independent reproduction easier for non-technical users and AI evaluators.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Meta-analysis and systematic reviews
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.