Search | VHL Regional Portal

A perceptual similarity space for speech based on self-supervised speech representations.

Chernyak, Bronya R; Bradlow, Ann R; Keshet, Joseph; Goldrick, Matthew.

J Acoust Soc Am ; 155(6): 3915-3929, 2024 Jun 01.

Article in English | MEDLINE | ID: mdl-38904539

ABSTRACT

Speech recognition by both humans and machines frequently fails in non-optimal yet common situations. For example, word recognition error rates for second-language (L2) speech can be high, especially under conditions involving background noise. At the same time, both human and machine speech recognition sometimes shows remarkable robustness against signal- and noise-related degradation. Which acoustic features of speech explain this substantial variation in intelligibility? Current approaches align speech to text to extract a small set of pre-defined spectro-temporal properties from specific sounds in particular words. However, variation in these properties leaves much cross-talker variation in intelligibility unexplained. We examine an alternative approach utilizing a perceptual similarity space acquired using self-supervised learning. This approach encodes distinctions between speech samples without requiring pre-defined acoustic features or speech-to-text alignment. We show that L2 English speech samples are less tightly clustered in the space than L1 samples reflecting variability in English proficiency among L2 talkers. Critically, distances in this similarity space are perceptually meaningful: L1 English listeners have lower recognition accuracy for L2 speakers whose speech is more distant in the space from L1 speech. These results indicate that perceptual similarity may form the basis for an entirely new speech and language analysis approach.

Subject(s)

Speech Acoustics , Speech Intelligibility , Speech Perception , Humans , Male , Female , Adult , Young Adult , Multilingualism , Recognition, Psychology , Noise

Automatic recognition of second language speech-in-noise.

Kim, Seung-Eun; Chernyak, Bronya R; Seleznova, Olga; Keshet, Joseph; Goldrick, Matthew; Bradlow, Ann R.

JASA Express Lett ; 4(2)2024 Feb 01.

Article in English | MEDLINE | ID: mdl-38350077

ABSTRACT

Measuring how well human listeners recognize speech under varying environmental conditions (speech intelligibility) is a challenge for theoretical, technological, and clinical approaches to speech communication. The current gold standard-human transcription-is time- and resource-intensive. Recent advances in automatic speech recognition (ASR) systems raise the possibility of automating intelligibility measurement. This study tested 4 state-of-the-art ASR systems with second language speech-in-noise and found that one, whisper, performed at or above human listener accuracy. However, the content of whisper's responses diverged substantially from human responses, especially at lower signal-to-noise ratios, suggesting both opportunities and limitations for ASR--based speech intelligibility modeling.

Subject(s)

Speech Perception , Humans , Speech Perception/physiology , Noise/adverse effects , Speech Intelligibility/physiology , Speech Recognition Software , Recognition, Psychology

ABSTRACT

Subject(s)

ABSTRACT

Subject(s)

SEND TO:

SELECTION OF CITATIONS

SEARCH DETAIL