i
Preliminary study DDH ()


Preliminary study

Screenshot%20from%20web-based%20interface%20for%20mutual%20human-machine%20learning%20phase%20of%20HMPL-C2-E1%20preliminary%20study.

Screenshot from web-based interface for mutual human-machine learning phase of HMPL-C2-E1 preliminary study.

Learner 1 (L1) - is a 5-year old – pre-school bilingual (90% German, 10% Slovak) daughter of the main author of this article

three HMPL-C2 exercise 1 (E1) sessions were executed on days 1, 3 and 5 of the study

each HMPL-C2-E1 session consisted of human-testing phase followed by a mutual human-machine learning phase

in each phase, sequences consisted of 5 repetitions of syllables started with occlusive labial consonant M or B and followed by the vowel A, E, I, O or U, thus yielding sequences from “MA MA MA MA MA” to “BU BU BU BU BU"

speech recordings collected during the learning phase subsequently provided input for the acoustic-model fine-tuning process

Evaluation & Results

Day 1

Day 3

Day 5

DeepSpeech_DE

0.96

0.84

0.64

KIds0

0.74

0.72

0.68

KIdsL1-1

0.69

0.78

0.44

KIdsL1-3

0.69

0.8

0.52

KIdsL1-5

0.69

0.74

0.48


sequences of five vowels resp. CV syllables which were displayed by DP were considered to provide the “reference”; output of the model yielded the hypotheses

data provided by L1 during three testing phases on days 1, 3 and 5 were evaluated by means of 5 different models

DeepSpeech_de = baseline model; KIds-0 : Deepspeech_de fine-tuned with kidsTALC; KIds-L1-1 : KIds-0 fine-tuned with data provided by L1 during day 1 learning phase;KIds-L1-3 : KIds-1 fine-tuned with data provided by L1 during day 3 learning phase;KIds-L1-5 : KIds-5 fine-tuned with data provided by L1 during day 5 learning phase

NOTE: WER-decrease between rows corresponds to increase of accuracy of the ASR model; WER-decrease between columns points to increase in L1's reading competence