Jev and adapted encoders on ToxicChat: operating-point comparison and incomplete budget evidence
Toxicity detectors must balance recall against false positives, while checkpoint comparisons and adaptation procedures require different evidence. We study 4,940 previously observed public ToxicChat examples: 2,724 have human annotations (348 toxic, 2,376 non-toxic), and 2,216 have automatic non-toxic labels. At operating points selected using 1,000 development labels, a recorded Jev API snapshot achieved 90.52% recall at a 3.24% false-positive rate (FPR), versus 35.63%–58.91% recall for three fixed encoder conditions. Only Jev met the specified absolute quality bounds on this mixed-label cohort. Post hoc analyses at unchanged thresholds found a 5.39% Jev FPR on the human-annotated subset and 120–196 toxic examples detected only by Jev in each encoder comparison, versus 5–12 detected only by the encoder. A separate adaptation campaign used budgets up to 3,105 benchmark training labels. Complete three-seed points for primary ToxicBERT achieved mean recalls of 50.77% at 64 labels and 59.77% at 256 labels; secondary mDeBERTa achieved 79.50% at 3,105 labels. No complete point established the fixed noninferiority and absolute-quality criteria. Observed primary N1024 checkpoints already fail the recall condition, while N3105 is unobserved, leaving the minimum qualifying tested budget indeterminate. Budgets count existing benchmark-label access, not new human annotation costs. Prior test observation, mixed label origins, differing implementations and limited replication restrict generalization. The study provides a conditional implementation comparison and a bounded account of incomplete adaptation evidence. This manuscript reports a conditional comparison on a previously observed public ToxicChat cohort and an incomplete label-budget adaptation campaign. It is a preprint and has not undergone formal human peer review. The manuscript reports numerical reproduction boundaries; independent acquisition and replay of the 94 remote learning-curve artifacts, checkpoint restoration and historical API replay remain unverified. Underlying raw data, model checkpoints and restricted remote artifacts are not included in this deposit.
Authors
- Takahiro Hagino
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-01
- DOI
- https://doi.org/10.5281/zenodo.23075657
- Primary Topic
- Hate Speech and Cyberbullying Detection
- Type
- preprint