A Human-Centered Web Technology for Reproducible Statistical and Machine-Learning Analysis of Heterogeneous Data

The analysis of heterogeneous tabular datasets remains challenging for non-specialist users because data preparation, statistical diagnostics, algorithm selection, model evaluation, and result interpretation are frequently distributed across separate software environments. This manuscript presents a reproducible web-based technology that integrates data ingestion, automated data-quality assessment, statistical validation, transparent rule-based analytical guidance, supervised and unsupervised machine learning, and interactive visual analytics within a unified workflow. The platform supports commonly used tabular formats and employs a modular, containerized architecture to support portability and reproducible analytical execution. Rather than autonomously selecting a supposedly optimal model, the recommendation mechanism identifies the analytical task, screens incompatible methods, detects semantic ambiguity, generates methodological warnings, and prioritizes compatible candidate models using explicit and inspectable rules. Final supervised-model selection is based on empirical cross-validation within the training data. The technology is evaluated through functional tests on heterogeneous health and financial datasets, predictive and clustering experiments, controlled recommendation-engine and schema-robustness experiments, computational benchmarks, and a usability study involving users with limited programming experience. In a COVID-19 dataset containing 95,839 records, the platform supported two-code and three-code patient-classification experiments, descriptive clustering, and data-quality analysis while transparently identifying class imbalance, predictor missingness, repeated predictor profiles, and performance limitations. A grouped-profile sensitivity analysis further showed that the principal classification results did not deteriorate when identical selected predictor configurations were prevented from occurring in both the training and held-out partitions. Controlled schema-robustness experiments correctly identified all prespecified unambiguous analytical tasks, detected all prespecified semantic ambiguities, rejected incompatible supervised configurations, and produced no silent acceptance of ambiguous or invalid cases. The recommender-specific evaluation also showed that the highest-ranked rule-based candidate does not necessarily coincide with the cross-validation-best model, reinforcing its role as a decision-support mechanism rather than an autonomous model-selection procedure. An exploratory laboratory usability evaluation with twelve non-specialist participants produced an average task-success rate of 95.3% and a mean System Usability Scale score of 86.7±5.2. These values correspond to aggregate results recorded during the original laboratory experiment. Because the participant-level observer logs and individual SUS responses were not retained in the available study materials, the usability results are interpreted descriptively and are not presented as population-level estimates. The results indicate that the proposed architecture can reduce technical barriers while preserving analytical traceability, methodological transparency, and user control. The platform therefore provides an extensible foundation for human-supervised, reproducible analysis of heterogeneous tabular data rather than an autonomous statistical or machine-learning decision-making system.

Authors

Institutions

Publication Details

Journal
Technologies
Published
2026-10-09
DOI
https://doi.org/10.3390/technologies14100648
Primary Topic
Scientific Computing and Data Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A Human-Centered Web Technology for Reproducible Statistical and Machine-Learning Analysis of Heterogeneous Data

E.-G. Espinosa–Martínez, Sergio Quezada–García, Mariano Vargas-Santiago, Ricardo Isaac Cázares Ramírez et al.
Technologies
Scientific Computing and Data Management
article

A Human-Centered Web Technology for Reproducible Statistical and Machine-Learning Analysis of Heterogeneous Data

E.-G. Espinosa–Martínez, Sergio Quezada–García, Mariano Vargas-Santiago, Ricardo Isaac Cázares Ramírez, Diana A. León-Velasco
article en

Abstract

The analysis of heterogeneous tabular datasets remains challenging for non-specialist users because data preparation, statistical diagnostics, algorithm selection, model evaluation, and result interpretation are frequently distributed across separate software environments. This manuscript presents a reproducible web-based technology that integrates data ingestion, automated data-quality assessment, statistical validation, transparent rule-based analytical guidance, supervised and unsupervised machine learning, and interactive visual analytics within a unified workflow. The platform supports commonly used tabular formats and employs a modular, containerized architecture to support portability and reproducible analytical execution. Rather than autonomously selecting a supposedly optimal model, the recommendation mechanism identifies the analytical task, screens incompatible methods, detects semantic ambiguity, generates methodological warnings, and prioritizes compatible candidate models using explicit and inspectable rules. Final supervised-model selection is based on empirical cross-validation within the training data. The technology is evaluated through functional tests on heterogeneous health and financial datasets, predictive and clustering experiments, controlled recommendation-engine and schema-robustness experiments, computational benchmarks, and a usability study involving users with limited programming experience. In a COVID-19 dataset containing 95,839 records, the platform supported two-code and three-code patient-classification experiments, descriptive clustering, and data-quality analysis while transparently identifying class imbalance, predictor missingness, repeated predictor profiles, and performance limitations. A grouped-profile sensitivity analysis further showed that the principal classification results did not deteriorate when identical selected predictor configurations were prevented from occurring in both the training and held-out partitions. Controlled schema-robustness experiments correctly identified all prespecified unambiguous analytical tasks, detected all prespecified semantic ambiguities, rejected incompatible supervised configurations, and produced no silent acceptance of ambiguous or invalid cases. The recommender-specific evaluation also showed that the highest-ranked rule-based candidate does not necessarily coincide with the cross-validation-best model, reinforcing its role as a decision-support mechanism rather than an autonomous model-selection procedure. An exploratory laboratory usability evaluation with twelve non-specialist participants produced an average task-success rate of 95.3% and a mean System Usability Scale score of 86.7±5.2. These values correspond to aggregate results recorded during the original laboratory experiment. Because the participant-level observer logs and individual SUS responses were not retained in the available study materials, the usability results are interpreted descriptively and are not presented as population-level estimates. The results indicate that the proposed architecture can reduce technical barriers while preserving analytical traceability, methodological transparency, and user control. The platform therefore provides an extensible foundation for human-supervised, reproducible analysis of heterogeneous tabular data rather than an autonomous statistical or machine-learning decision-making system.

TechnologiesVol. 14(10)
Universidad Autónoma de la Ciudad de México (MX), Universidad Autónoma Metropolitana (MX), Secretaría de Ciencia Tecnología e Innovación (MX), Secretaría de Ciencia, Humanidades, Tecnología e Innovación (MX), Universidad del Valle de México (MX), Universidad Nacional Autónoma de México (MX)
Openalex Percentile: Top 7%
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.