dataSDA: datasets and basic statistics for symbolic data analysis in R

Traditional datasets typically represent each variable with a single value per observation. However, as data volume and complexity continue to grow, representing variables through high-level descriptors such as intervals, histograms, and probability distributions, collectively referred to as symbolic data, has become increasingly valuable. These enriched representations retain information about distributional structure and variability, thereby enhancing both analytical depth and interpretability. This paper introduces dataSDA, an R package developed to curate symbolic datasets across diverse research domains and to support their reading, writing, conversion, and summarization. Building upon the frameworks of RSDA and HistDAWass, dataSDA extends their functionality by providing unified format conversion with automatic detection, aggregation of conventional (single-valued) data into symbolic form, and functions for computing interval distances, similarity measures, and descriptive statistics. The package currently hosts 114 benchmark datasets. A subset of these is used to illustrate clustering, classification, and regression analyses, as well as exploratory data analysis and visualization with the ggInterval package. By integrating ready-to-use symbolic datasets with tools for data format transformation and descriptive statistics, dataSDA aims to serve as a comprehensive resource for symbolic data collection and analysis. The package promotes accessibility, transparency, and reproducibility in symbolic data research and is freely available on CRAN.

Authors

Institutions

Publication Details

Journal
Journal of Applied Statistics
Published
2026-09-21
DOI
https://doi.org/10.1080/02664763.2026.2730249
Primary Topic
Advanced Statistical Modeling Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

dataSDA: datasets and basic statistics for symbolic data analysis in R

Han‐Ming Wu, Chun‐Houh Chen, Po-Wei Chen
Journal of Applied Statistics
Advanced Statistical Modeling Techniques
article

dataSDA: datasets and basic statistics for symbolic data analysis in R

Han‐Ming Wu, Chun‐Houh Chen, Po-Wei Chen
article en

Abstract

Traditional datasets typically represent each variable with a single value per observation. However, as data volume and complexity continue to grow, representing variables through high-level descriptors such as intervals, histograms, and probability distributions, collectively referred to as symbolic data, has become increasingly valuable. These enriched representations retain information about distributional structure and variability, thereby enhancing both analytical depth and interpretability. This paper introduces dataSDA, an R package developed to curate symbolic datasets across diverse research domains and to support their reading, writing, conversion, and summarization. Building upon the frameworks of RSDA and HistDAWass, dataSDA extends their functionality by providing unified format conversion with automatic detection, aggregation of conventional (single-valued) data into symbolic form, and functions for computing interval distances, similarity measures, and descriptive statistics. The package currently hosts 114 benchmark datasets. A subset of these is used to illustrate clustering, classification, and regression analyses, as well as exploratory data analysis and visualization with the ggInterval package. By integrating ready-to-use symbolic datasets with tools for data format transformation and descriptive statistics, dataSDA aims to serve as a comprehensive resource for symbolic data collection and analysis. The package promotes accessibility, transparency, and reproducibility in symbolic data research and is freely available on CRAN.

Journal of Applied Statistics
Academia Sinica (TW), National Chengchi University (TW), National Taipei University (TW)
Quality Education
Openalex Percentile: Top 9%
Advanced Statistical Modeling Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.