Machine learning and network-based classification approach in opium and tobacco users: results from Fasa PERSIAN cohort study
Opium and tobacco are widely used in parts of the Middle East and South Asia, yet their individual and combined effects on routine hematologic and biochemical markers are not fully understood. We examined how tobacco smoking (TS), opium use (OU), and their combination (OUTS) relate to laboratory parameters and the ability to distinguish user groups from never-users (Normal). We analyzed 3497 participants from the baseline Fasa PERSIAN Cohort, classified into Normal, TS, OU, and OUTS groups. Group-specific correlation networks and community detection were used to explore network structure in clinical laboratory parameters. Multiple machine learning models were trained for classification using standardized features and evaluated by stratified 5-fold cross-validation. An exploratory OWA ensemble was additionally evaluated using validation-based selection and an independent test set. Performance was assessed using accuracy, sensitivity, specificity, ROC AUC, and balanced accuracy. Network topology differed across all four groups, with the greatest disruption in OUTS, particularly in lipid–liver function correlations and central nodes such as HDL.C. Differentiation between OUTS and TS was limited. In the secondary single-model ML analysis, the normal group was more effectively distinguished from user groups, with maximum AUCs of 0.81 (TS), 0.76 (OU), and 0.87 (OUTS). For OUTS versus OU, the primary model achieved 88% acceptable sensitivity and 35% low specificity. Across the three OWA configurations for OUTS versus TS, accuracy ranged from 65.6 to 66.9%, AUC from 0.586 to 0.656, sensitivity from 93.3 to 97.1%, and specificity from 3.8 to 15.1%. This study combined machine learning and network base approach to unravel the complex effects of opium and tobacco on blood parameters. The secondary Normal-versus-user analyses showed greater discrimination than the substance-user comparisons. Single-model ML showed lower discrimination for the substance-user comparisons, particularly OUTS versus TS, while the exploratory OWA analysis of OUTS versus TS also showed limited class-balanced discrimination
Authors
- Fatemeh Vafaee Sharbaf
- Azizallah Dehghan (ORCID: https://orcid.org/0000-0002-7345-0796)
- Mojtaba Farjam (ORCID: https://orcid.org/0000-0003-4826-2846)
- Samaneh Toutounchian (ORCID: https://orcid.org/0000-0002-1975-2697)
- Kaveh Kavousi (ORCID: https://orcid.org/0000-0002-1906-3912)
- Amirreza Dehghanian (ORCID: https://orcid.org/0000-0003-0360-7468)
- Zahra Salehi (ORCID: https://orcid.org/0000-0002-0839-2729)
- Mohammad Mehdi Naghizadeh (ORCID: https://orcid.org/0000-0001-5562-103X)
- Kiarash Zare
- Marzieh Gholami
- Alireza Morovat
- Mohamad Reza Zabihi
Institutions
- Shiraz University of Medical Sciences (IR)
- University of Tehran (IR)
- Fasa University of Medical Sciences (IR)
- Tehran University of Medical Sciences (IR)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1038/s41598-026-74430-6
- Primary Topic
- Artificial Intelligence in Healthcare
- Type
- article
- Field-Weighted Citation Impact
- 0.00