Improving big data analytics ecosystems using ad-hoc parallel file systems

Data processing in different technology areas, such as Artificial Intelligence or Big Data, has recently challenged High-Performance Computing (HPC). This has led to innovation in processing and managing these huge volumes of data. Numerous systems have sought to address this big-data issue. Parallel file systems, such as Expand, use techniques such as file partitioning or replication to provide highly available, high-performance storage systems for HPC environments. In addition, several frameworks have been deployed for use with parallel file systems in Big Data Analytics (BDA) environments to reduce bottlenecks caused by massive input/output (I/O) operations. This article presents a new solution for such BDA ecosystems using Apache Spark and Expand. Expand is a parallel and distributed file system designed by the ARCOS research group that can be used as a parallel ad-hoc file system to alleviate I/O bottlenecks arising in traditional parallel file systems. Using Apache Spark in conjunction with the Expand file system enables you to leverage the benefits of both platforms. Evaluations presented in this paper compare Expand with other file systems (Lustre and HDFS) and demonstrate the advantages of an ad-hoc parallel file system for BDA applications. By using the TeraSort benchmark in the evaluation, Expand was found to deliver a performance that is up to 4.5 times higher than Lustre and 2.5 times higher than HDFS. This benchmark is widely used for evaluating BDA applications.

Authors

Institutions

Publication Details

Journal
Journal Of Big Data
Published
2026-09-22
DOI
https://doi.org/10.1186/s40537-026-01559-6
Primary Topic
Advanced Data Storage Technologies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Improving big data analytics ecosystems using ad-hoc parallel file systems

Félix Garcı́a-Carballeira, Diego Camarmas-Alonso, Alejandro Calderón, Dario Muñoz-Muñoz et al.
Journal Of Big Data
Advanced Data Storage Technologies
article

Improving big data analytics ecosystems using ad-hoc parallel file systems

Félix Garcı́a-Carballeira, Diego Camarmas-Alonso, Alejandro Calderón, Dario Muñoz-Muñoz, Gabriel Sotodosos-Morales, Jesus Carretero
article en

Abstract

Data processing in different technology areas, such as Artificial Intelligence or Big Data, has recently challenged High-Performance Computing (HPC). This has led to innovation in processing and managing these huge volumes of data. Numerous systems have sought to address this big-data issue. Parallel file systems, such as Expand, use techniques such as file partitioning or replication to provide highly available, high-performance storage systems for HPC environments. In addition, several frameworks have been deployed for use with parallel file systems in Big Data Analytics (BDA) environments to reduce bottlenecks caused by massive input/output (I/O) operations. This article presents a new solution for such BDA ecosystems using Apache Spark and Expand. Expand is a parallel and distributed file system designed by the ARCOS research group that can be used as a parallel ad-hoc file system to alleviate I/O bottlenecks arising in traditional parallel file systems. Using Apache Spark in conjunction with the Expand file system enables you to leverage the benefits of both platforms. Evaluations presented in this paper compare Expand with other file systems (Lustre and HDFS) and demonstrate the advantages of an ad-hoc parallel file system for BDA applications. By using the TeraSort benchmark in the evaluation, Expand was found to deliver a performance that is up to 4.5 times higher than Lustre and 2.5 times higher than HDFS. This benchmark is widely used for evaluating BDA applications.

Journal Of Big Data
Universidad Carlos III de Madrid (ES)
Industry, innovation and infrastructure
Openalex Percentile: Top 9%
Advanced Data Storage Technologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Improving big data analytics ecosystems using ad-hoc parallel file systems — Félix Garcı́a-Carballeira, Diego Camarmas-Alonso, et al. · Journal Of Big Data (2026) | TGRS Research Map | TGRS