Sino-US-DrugQA: A benchmark for evaluating large language models in cross-jurisdictional pharmaceutical regulation
Cross-jurisdictional pharmaceutical compliance requires comparison of regulatory requirements across administrative systems such as the US Food and Drug Administration and China's National Medical Products Administration. Although large language models (LLMs) are increasingly explored for healthcare and regulatory applications, their performance in cross-jurisdictional pharmaceutical regulation has not been systematically evaluated using a dedicated benchmark. We introduce Sino-US-DrugQA, a bilingual multiple-choice benchmark covering Monolingual, Comparative, and Parallel regulatory question–answer tasks. The 11,871 candidate items underwent deterministic structural validation and full-dataset semantic quality screening, followed by risk-stratified independent review of 1,432 items by two regulatory experts. The final release comprised 11,444 items, including 10,122 classified as pass and 1,322 as borderline. Among 500 items sampled from the semantic screen-negative population, 18 were subsequently classified as material errors, corresponding to an observed residual material-error proportion of 3.60% (Wilson 95% confidence interval, 2.29%–5.62%). Four representative LLMs—gpt-5.6-terra, gemini-3.6-flash, deepseek-v4-flash, and qwen-3.5-max—were evaluated under a standardized zero-shot protocol. Overall accuracy ranged from 83.21% to 86.43%. Comparative accuracy was consistently lower than Monolingual accuracy, with absolute differences of 4.42–8.98 percentage points across models. The two highest-scoring models, GPT and Gemini, did not differ significantly after adjustment for multiple comparisons. These results indicate that explicit comparison across non-equivalent regulatory systems remains more challenging than single-jurisdiction question answering. Sino-US-DrugQA provides a validated and reproducible resource for evaluating bilingual regulatory reasoning. The findings support further investigation of expert-supervised decision-support workflows rather than autonomous regulatory interpretation. The stable dataset release and evaluation resources are available at https://github.com/DodgeLU/Sino-US-DrugQA .
Authors
- Zhen Chen (ORCID: https://orcid.org/0000-0002-1810-0195)
- Xuejing Fu
- Wentao Lu
Institutions
- Beijing Academy of Artificial Intelligence (CN)
- Shanghai Drug Administration (CN)
- Shanghai Municipal People's Government (CN)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1371/journal.pone.0343858
- Primary Topic
- Pharmacovigilance and Adverse Drug Reactions
- Type
- article
- Field-Weighted Citation Impact
- 0.00