i
Diese Präsentation ist für die Ansicht im Querformat gestaltet.
Mozilla's DeepSpeech ASR Architecture
reading is essentially a process of translation of textual sequences into their phonetic representations
spoken word thus play a fundamental role in reading acquisition
highly accurate automatic speech recognition (ASR ) systems exist for many languages but they are still strongly biased towards accurate processing of adult voices
HOWEVER: in reading acquisition or reading fostering scenarios one deals with speakers whoseutterances of sequences-to-be-read exhibit peculiar characteristics
majority of those who learn how to read are children
children are physiologically (differences in size and anatomy of vocal tract; teeth change) and cognitively different from adults
children voices are different from adult voices (e.g. Fundamental frequency F of male voice =~ 112.0 Hz; F(female voice) =~ 195.8 Hz; F(boy voice) =~ 250.0 Hz; F(girl voice) =~ 244.0 Hz
datasets for children’s speech that are publicly available are quite scarce
kidsTALC (Rumberg et al, 2022) authors report 26.2 % word-error-rate (WER) of typically developing monolingual German children