Applied Machine Learning for Cross-Substance Analysis: Psychometric and Demographic Correlates of Consumption Patterns
This study provides an end-to-end, reproducible empirical machine learning benchmark evaluating the relationship between psychometric traits and self-reported substance consumption on the benchmark UCI Drug Consumption dataset. Operating under rigorous data hygiene protocols, we systematically eliminate fictitious drug (Semeron) overclaimers (N=8), yielding an analytically verified cohort of N=1,877 respondents characterized across 10 continuous demographic and psychometric dimensions (NEO-FFI-R Five-Factor Model, BIS-11 Impulsivity, and ImpSS Sensation Seeking). The investigation establishes two decoupled yet mutually reinforcing machine learning paradigms: (1) Unsupervised Behavioral Segmentation via vector-optimized K-Means and soft Fuzzy C-Means (FCM) from mathematical first principles, identifying a four-cluster topology with exploratory non-parametric Kruskal-Wallis testing; and (2) Cross-Substance Supervised Tournament benchmarking six predictive model configurations across four illicit substance classes (Cannabinoids, CNS Stimulants, Psychedelics, and Depressants/Anxiolytics). Generalization is evaluated across 5-fold outer cross-validation and frozen holdouts with strict scaling isolation, 95% Wald confidence intervals via inverted Hessian covariance, Benjamini-Hochberg False Discovery Rate control (q=0.05), and empirical probability calibration diagnostics. Code repository: https://github.com/rishindra-mateti-tech/applied-machine-learning-cross-substance-analysis
Authors
- Rishindra Mateti
Institutions
- Wright State University (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23166994
- Primary Topic
- Data Mining and Machine Learning Applications
- Type
- preprint