Big Data Management in the Modern Era: Application of Statistical Tools for Large-Scale Data Analysis
This paper explores the critical relationship between large-scale data management infrastructure and modern statistical methodologies in the modern era. While high-performance distributed computing frameworks such as Apache Hadoop, Apache Spark, and cloud-based object storage (AWS S3, Azure Blob, GCS) address the technical challenges of handling immense volume, velocity, and variety, computational power alone cannot derive actionable insights. To bridge this gap, this study outlines a structured four-stage analytical framework—spanning Data Ingestion, Storage & Automated Preprocessing, Parallelized Statistical Analytics, and Insight Generation. It details the adaptation of advanced statistical techniques (such as Principal Component Analysis, Stochastic Gradient Descent, Time-Series Analysis, and Reservoir Sampling) for distributed cluster environments. Furthermore, the paper discusses real-world industrial applications across finance, healthcare, e-commerce, and smart cities, while addressing critical operational challenges including data quality, network latency, and privacy compliance (GDPR/Differential Privacy).
Authors
- Saikat Deb (ORCID: https://orcid.org/0009-0000-3358-6982)
Institutions
- Sylhet Agricultural University (BD)
- Leading University (BD)
- Sylhet International University (BD)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22963403
- Primary Topic
- Cloud Computing and Resource Management
- Type
- preprint