Jev and adapted encoders on ToxicChat: operating-point comparison and incomplete budget evidence

Toxicity detectors must balance recall against false positives, while checkpoint comparisons and adaptation procedures require different evidence. We study 4,940 previously observed public ToxicChat examples: 2,724 have human annotations (348 toxic, 2,376 non-toxic), and 2,216 have automatic non-toxic labels. At operating points selected using 1,000 development labels, a recorded Jev API snapshot achieved 90.52% recall at a 3.24% false-positive rate (FPR), versus 35.63%–58.91% recall for three fixed encoder conditions. Only Jev met the specified absolute quality bounds on this mixed-label cohort. Post hoc analyses at unchanged thresholds found a 5.39% Jev FPR on the human-annotated subset and 120–196 toxic examples detected only by Jev in each encoder comparison, versus 5–12 detected only by the encoder. A separate adaptation campaign used budgets up to 3,105 benchmark training labels. Complete three-seed points for primary ToxicBERT achieved mean recalls of 50.77% at 64 labels and 59.77% at 256 labels; secondary mDeBERTa achieved 79.50% at 3,105 labels. No complete point established the fixed noninferiority and absolute-quality criteria. Observed primary N1024 checkpoints already fail the recall condition, while N3105 is unobserved, leaving the minimum qualifying tested budget indeterminate. Budgets count existing benchmark-label access, not new human annotation costs. Prior test observation, mixed label origins, differing implementations and limited replication restrict generalization. The study provides a conditional implementation comparison and a bounded account of incomplete adaptation evidence. This manuscript reports a conditional comparison on a previously observed public ToxicChat cohort and an incomplete label-budget adaptation campaign. It is a preprint and has not undergone formal human peer review. The manuscript reports numerical reproduction boundaries; independent acquisition and replay of the 94 remote learning-curve artifacts, checkpoint restoration and historical API replay remain unverified. Underlying raw data, model checkpoints and restricted remote artifacts are not included in this deposit.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-01
DOI
https://doi.org/10.5281/zenodo.23075657
Primary Topic
Hate Speech and Cyberbullying Detection
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Jev and adapted encoders on ToxicChat: operating-point comparison and incomplete budget evidence

Takahiro Hagino
Zenodo (CERN European Organization for Nuclear Research)
Hate Speech and Cyberbullying Detection
preprint

Jev and adapted encoders on ToxicChat: operating-point comparison and incomplete budget evidence

Takahiro Hagino
preprint en

Abstract

Toxicity detectors must balance recall against false positives, while checkpoint comparisons and adaptation procedures require different evidence. We study 4,940 previously observed public ToxicChat examples: 2,724 have human annotations (348 toxic, 2,376 non-toxic), and 2,216 have automatic non-toxic labels. At operating points selected using 1,000 development labels, a recorded Jev API snapshot achieved 90.52% recall at a 3.24% false-positive rate (FPR), versus 35.63%–58.91% recall for three fixed encoder conditions. Only Jev met the specified absolute quality bounds on this mixed-label cohort. Post hoc analyses at unchanged thresholds found a 5.39% Jev FPR on the human-annotated subset and 120–196 toxic examples detected only by Jev in each encoder comparison, versus 5–12 detected only by the encoder. A separate adaptation campaign used budgets up to 3,105 benchmark training labels. Complete three-seed points for primary ToxicBERT achieved mean recalls of 50.77% at 64 labels and 59.77% at 256 labels; secondary mDeBERTa achieved 79.50% at 3,105 labels. No complete point established the fixed noninferiority and absolute-quality criteria. Observed primary N1024 checkpoints already fail the recall condition, while N3105 is unobserved, leaving the minimum qualifying tested budget indeterminate. Budgets count existing benchmark-label access, not new human annotation costs. Prior test observation, mixed label origins, differing implementations and limited replication restrict generalization. The study provides a conditional implementation comparison and a bounded account of incomplete adaptation evidence. This manuscript reports a conditional comparison on a previously observed public ToxicChat cohort and an incomplete label-budget adaptation campaign. It is a preprint and has not undergone formal human peer review. The manuscript reports numerical reproduction boundaries; independent acquisition and replay of the 94 remote learning-curve artifacts, checkpoint restoration and historical API replay remain unverified. Underlying raw data, model checkpoints and restricted remote artifacts are not included in this deposit.

Zenodo (CERN European Organization for Nuclear Research)
Hate Speech and Cyberbullying Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Jev and adapted encoders on ToxicChat: operating-point comparison and incomplete budget evidence — Takahiro Hagino · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS