Deep Sound Synthesis Matched to Brain Activity Recapitulates Preferential Responses to Speech and Music

L Lidongsheng Xing (邢立冬生) E Elia Formisano L Lars Riecke

Abstract

The human auditory system extracts meaning from sounds in the environment by transforming acoustic input signals into semantic categories, such as speech and music. Although distinct acoustic features give rise to these categorical percepts and to preferential responses in spatially segregated regions in the auditory cortex, the nature of the internal representations underlying this transformation remains poorly understood. Here, we combined neuroimaging, a deep neural network (DNN), brain-based sound synthesis, and psychophysical testing in human participants of either sex to investigate the internal sound features encoded in speech- and music-selective regions of the auditory cortex and their functional role in sound categorization. We found that sounds synthesized from cortical activity patterns—though acoustically dissimilar to natural speech and music sounds—nonetheless elicited similar categorical cortical and behavioral responses. These results suggest that the auditory cortex relies on internal, abstracted representations of category structure that are not reducible to the natural acoustic properties of speech and music. Our findings provide new insights into intermediate sound features, as captured by DNNs that may support categorization in the human auditory system.

Article Details

Volume / Issue Vol. 46, Issue 9
Published March 04, 2026
Pages e1651252025
ISSN 0270-6474
Publisher Society for Neuroscience

Journal Info

Journal of Neuroscience

Society for Neuroscience

ISSN: 0270-6474 Life Sciences

Authors (3)

L

Lidongsheng Xing (邢立冬生)

E

Elia Formisano

L

Lars Riecke