Beyond Keyword Filters: Calibrated Monte-Carlo Risk Gating for Safe Multilingual Colorectal-Cancer LLM Dialogue
Large language models are increasingly consulted by cancer patients, and a single unsafe answer about chemotherapy dosing, opioid use or self-harm can cause real harm. This paper introduces a calibrated Monte-Carlo risk gate that treats colorectal-cancer dialogue safety as a selective-prediction problem, estimating the risk of a user turn from a bootstrap ensemble over multilingual sentence representations, calibrating it with Platt scaling and deciding at a single threshold between an informative answer and referral to a clinician. Evaluated on 450 oncologist-approved prompts in English, Turkish and Spanish under a scenario-level split, the gate reaches a guardrail F1 of 0.961, blocks 98.2 percent of harmful prompts and refuses 10.8 percent of legitimate questions, while the keyword, regular-expression and fuzzy layers that dominate current practice reach at most 0.034 and fire on five of 450 prompts, a separation that holds at a corrected q of 0.0003 and survives Bonferroni correction. Ablation locates the mechanism, with the semantic representation carrying the discriminative signal, Platt scaling lowering the expected calibration error from 0.216 to 0.069, and the ensemble predicting its own errors at an AUROC of 0.856 and removing them entirely at 50 percent coverage. Calibrated selective prediction over semantic representations makes multilingual medical-dialogue safety measurable, tunable to an explicit operating point, and consistent across the three languages tested.
Authors
- Kerem Gencer (ORCID: https://orcid.org/0000-0002-2914-1056)
- Abdurrahim Kızılay
Institutions
- Afyon Kocatepe University (TR)
Publication Details
- Journal
- Big Data and Cognitive Computing
- Published
- 2026-09-14
- DOI
- https://doi.org/10.3390/bdcc10090317
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00