A Chemical Foundation Model Benchmark for PFAS Bioactivity Prediction: Multi-Task Fine-Tuning Beats Fingerprint Baselines, PFAS Domain Pretraining Does Not Help

Abstract Per- and polyfluoroalkyl substances (PFAS) comprise more than 21,000 chemicals, yet fewer than 1.3% have any experimental bioactivity measurement, making structure-based hazard prediction the gating problem for prioritizing the untested majority. Existing PFAS models are single-task and rarely evaluated under leakage-aware cross-validation. We present the first multitask chemical foundation-model benchmark for PFAS bioactivity, fine-tuning off-the-shelf MolFormer-XL with 156 ToxCast classification heads under Tanimoto-Butina cluster-out cross-validation, which is appropriate because 89% of PFAS have empty Murcko scaffolds. The model reaches macro AUROC 0.764, beating a Morgan-fingerprint XGBoost baseline by +0.18 and winning on 155 of 156 end points. That baseline is, however, partly degenerate: binary Morgan fingerprints map a quarter of the compounds onto identical vectors, and sweeping eight 1D/2D/3D representations on the same folds raises the best classical model to 0.688, so the margin properly attributable to the foundation model is +0.076. A matched factorial ablation shows that backbone fine-tuning and multitask learning are synergistic: multitask learning helps only once the backbone is fine-tuned (+0.08). This interaction is backbone- and task-dependent: it reproduces on Tox21 for two backbones but is absent for ChemBERTa-2 on PFAS, and label-free encoder diagnostics show this is not because ChemBERTa-2 represents PFAS poorly, so static representation quality does not predict postfine-tuning performance. PFAS-specific domain-adaptation pretraining gives no detectable improvement (Δ = – 0.01 to – 0.03). Mondrian-by-class conformal calibration improves rare-active-class coverage from 0.36 to 0.78, and an applicability-domain analysis shows only 26% of untested PFAS lie close enough to the training set for the model to be applied without caveat. The representation also transfers to soil-water partitioning (Kd) regression, where it holds at Spearman ρ = 0.75 even on PFAS chemistries absent from training while the fingerprint baselines collapse.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Information and Modeling
Published
2026-09-24
DOI
https://doi.org/10.1021/acs.jcim.6c02202
Primary Topic
Per- and polyfluoroalkyl substances research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Chemical Foundation Model Benchmark for PFAS Bioactivity Prediction: Multi-Task Fine-Tuning Beats Fingerprint Baselines, PFAS Domain Pretraining Does Not Help

Yue Zhu, Qingyang Liu
Journal of Chemical Information and Modeling
Per- and polyfluoroalkyl substances research
article

A Chemical Foundation Model Benchmark for PFAS Bioactivity Prediction: Multi-Task Fine-Tuning Beats Fingerprint Baselines, PFAS Domain Pretraining Does Not Help

Yue Zhu, Qingyang Liu
article en

Abstract

Abstract Per- and polyfluoroalkyl substances (PFAS) comprise more than 21,000 chemicals, yet fewer than 1.3% have any experimental bioactivity measurement, making structure-based hazard prediction the gating problem for prioritizing the untested majority. Existing PFAS models are single-task and rarely evaluated under leakage-aware cross-validation. We present the first multitask chemical foundation-model benchmark for PFAS bioactivity, fine-tuning off-the-shelf MolFormer-XL with 156 ToxCast classification heads under Tanimoto-Butina cluster-out cross-validation, which is appropriate because 89% of PFAS have empty Murcko scaffolds. The model reaches macro AUROC 0.764, beating a Morgan-fingerprint XGBoost baseline by +0.18 and winning on 155 of 156 end points. That baseline is, however, partly degenerate: binary Morgan fingerprints map a quarter of the compounds onto identical vectors, and sweeping eight 1D/2D/3D representations on the same folds raises the best classical model to 0.688, so the margin properly attributable to the foundation model is +0.076. A matched factorial ablation shows that backbone fine-tuning and multitask learning are synergistic: multitask learning helps only once the backbone is fine-tuned (+0.08). This interaction is backbone- and task-dependent: it reproduces on Tox21 for two backbones but is absent for ChemBERTa-2 on PFAS, and label-free encoder diagnostics show this is not because ChemBERTa-2 represents PFAS poorly, so static representation quality does not predict postfine-tuning performance. PFAS-specific domain-adaptation pretraining gives no detectable improvement (Δ = – 0.01 to – 0.03). Mondrian-by-class conformal calibration improves rare-active-class coverage from 0.36 to 0.78, and an applicability-domain analysis shows only 26% of untested PFAS lie close enough to the training set for the model to be applied without caveat. The representation also transfers to soil-water partitioning (Kd) regression, where it holds at Spearman ρ = 0.75 even on PFAS chemistries absent from training while the fingerprint baselines collapse.

Journal of Chemical Information and Modeling
Northeastern University (US), Georgia Institute of Technology (US)
Openalex Percentile: Top 19%
Per- and polyfluoroalkyl substances research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.