i
Reading Acquisition and Automatic Speech Recognition DDH ()


Reading Acquisition and Automatic Speech Recognition

Mozilla's%20DeepSpeech%20ASR%20Architecture

Mozilla's DeepSpeech ASR Architecture

reading is essentially a process of translation of textual sequences into their phonetic representations

spoken word thus play a fundamental role in reading acquisition

highly accurate automatic speech recognition (ASR ) systems exist for many languages but they are still strongly biased towards accurate processing of adult voices

HOWEVER: in reading acquisition or reading fostering scenarios one deals with speakers whoseutterances of sequences-to-be-read exhibit peculiar characteristics

Child Speech Recognition

majority of those who learn how to read are children

children are physiologically (differences in size  and anatomy of vocal tract; teeth change) and cognitively different from adults

children voices are different from adult voices (e.g. Fundamental frequency F of male voice =~ 112.0 Hz; F(female voice) =~ 195.8 Hz; F(boy voice) =~ 250.0 Hz; F(girl voice) =~ 244.0 Hz

datasets for children’s speech that are publicly available are quite scarce

kidsTALC (Rumberg et al, 2022) authors report 26.2 % word-error-rate (WER) of typically developing monolingual German children