ECHO Corpus: Cross-cultural database of 15000 nonverbal vocalisations from speakers of 22 languages

Research on human non-verbal vocalisations such as screams, cries and laughter is rapidly expanding, yet most studies remain restricted to single cultures and heavily overrepresent English speakers. Here, we introduce the ECHO corpus – a large-scale cross-cultural database of non-verbal vocalisations produced by nearly 1050 vocalisers aged 18-84 years from 52 countries, including native speakers of 22 different languages spanning 10 language families. The corpus comprises over 15000 volitional vocalizations expressing sixteen affective states, from pain, aggression and alarm to amusement, relief and sensual pleasure, together with baseline speech from the same speakers. We provide extensive metadata on vocalisers (e.g., age, sex, countries of birth and residence, spoken languages, and media usage) and validate the dataset through acoustic quality assessment, analyses of fundamental frequency, and cross-cultural perception experiments demonstrating high recognition of vocalisation contexts. This first-of-its-kind resource provides a representative foundation for investigating how humans vocally encode and decode emotions and motivations across diverse cultures, revealing the mechanisms and functions of human vocal communication. The ECHO corpus is openly available and his content is described in the README file.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22754937
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ECHO Corpus: Cross-cultural database of 15000 nonverbal vocalisations from speakers of 22 languages

Katarzyna Pisanski, Hiroki Koda, Théophile Turco, David Reby et al.
Zenodo (CERN European Organization for Nuclear Research)
Emotion and Mood Recognition
article

ECHO Corpus: Cross-cultural database of 15000 nonverbal vocalisations from speakers of 22 languages

Katarzyna Pisanski, Hiroki Koda, Théophile Turco, David Reby, Aitana García Arasco, Isao Tokuda, Uara Isa
article en

Abstract

Research on human non-verbal vocalisations such as screams, cries and laughter is rapidly expanding, yet most studies remain restricted to single cultures and heavily overrepresent English speakers. Here, we introduce the ECHO corpus – a large-scale cross-cultural database of non-verbal vocalisations produced by nearly 1050 vocalisers aged 18-84 years from 52 countries, including native speakers of 22 different languages spanning 10 language families. The corpus comprises over 15000 volitional vocalizations expressing sixteen affective states, from pain, aggression and alarm to amusement, relief and sensual pleasure, together with baseline speech from the same speakers. We provide extensive metadata on vocalisers (e.g., age, sex, countries of birth and residence, spoken languages, and media usage) and validate the dataset through acoustic quality assessment, analyses of fundamental frequency, and cross-cultural perception experiments demonstrating high recognition of vocalisation contexts. This first-of-its-kind resource provides a representative foundation for investigating how humans vocally encode and decode emotions and motivations across diverse cultures, revealing the mechanisms and functions of human vocal communication. The ECHO corpus is openly available and his content is described in the README file.

Zenodo (CERN European Organization for Nuclear Research)
Centre National de la Recherche Scientifique (FR), Ritsumeikan University (JP), Tokyo University of the Arts (JP), Institut Universitaire de France (FR), University of Wrocław (PL), Laboratoire Dynamique du Langage (FR), Japan Graduate School of Education University (JP), Institute of Psychology (PL), The University of Tokyo (JP), Université Jean Monnet (FR)
Quality Education
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.