Synchronous Activity in the Auditory Forebrain Represents Temporal Landmarks in Vocal Communication Signals
Precise temporal segmentation of sound is critical for identifying and interpreting complex acoustic signals such as human speech (Pisoni and Luce, 1987) and bird song (Gentner, 2008). Using the zebra finch model system, we investigated the neural coding of the songs' temporal features. We recorded neural activity from avian primary and secondary auditory cortex-like brain areas using electrode arrays in awake birds (4 males and 3 females) exposed to conspecific songs. By applying a deep-learning approach with bidirectional long-short-term memory networks (Daly, 2023), we demonstrate that not only the amplitude envelope of song stimuli but also discrete, time-locked envelope landmarks can be decoded from multi-channel neuronal responses to songs. The times of envelope peak amplitude and peak slope were decoded with average accuracies of 40% and maximum accuracies of 70% at 20-ms time scales. The decoding of these landmarks enabled the segmentation of song into syllables with maximum accuracy of 90% at 10-ms time scales. Decoding accuracy for temporal events was greater in the primary auditory than in the secondary auditory area, suggesting a serial processing of the segmentation and identification operation. These findings reveal that neural entrainment to salient temporal events in vocal communication signals is a ubiquitous, specialized feature of the vertebrate primary auditory forebrain (Mesgarani et al., 2009; Heelan et al., 2019) providing a mechanism for segmenting continuous acoustic streams into behaviorally relevant units for higher-order processing. Significance Statement To decipher speech or vocal communication signals, humans and animals must segment continuous sound streams into identifiable syllables or call-types. We found that amplitude envelope of songs and the time of salient acoustic landmarks can be decoded with high-accuracy from the activity of neurons in the auditory cortex analog of songbirds. These landmarks can be used to segment the song into its component syllables. We found higher fidelity for the representation of the envelope of the sound in primary relative to secondary auditory regions. We postulate that this functional organization corresponds to the order of auditory computations where the segmentation task precedes the identification task.
Authors
- Hermina Robotka (ORCID: https://orcid.org/0000-0002-3523-032X)
- Frédéric E. Theunissen (ORCID: https://orcid.org/0000-0001-8980-0650)
- Amirmasoud Ahmadi (ORCID: https://orcid.org/0000-0003-4895-2150)
- Manfred Gahr
Publication Details
- Journal
- Journal of Neuroscience
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1523/jneurosci.2108-25.2026
- Primary Topic
- Animal Vocal Communication and Behavior
- Type
- article
- Field-Weighted Citation Impact
- 0.00