i
Diese Präsentation ist für die Ansicht im Querformat gestaltet.
Screenshot from web-based interface for mutual human-machine learning phase of HMPL-C2-E1 preliminary study.
Learner 1 (L1) - is a 5-year old – pre-school bilingual (90% German, 10% Slovak) daughter of the main author of this article
three HMPL-C2 exercise 1 (E1) sessions were executed on days 1, 3 and 5 of the study
each HMPL-C2-E1 session consisted of human-testing phase followed by a mutual human-machine learning phase
in each phase, sequences consisted of 5 repetitions of syllables started with occlusive labial consonant M or B and followed by the vowel A, E, I, O or U, thus yielding sequences from “MA MA MA MA MA” to “BU BU BU BU BU"
speech recordings collected during the learning phase subsequently provided input for the acoustic-model fine-tuning process
|
Day 1 |
Day 3 |
Day 5 |
DeepSpeech_DE |
0.96 |
0.84 |
0.64 |
KIds0 |
0.74 |
0.72 |
0.68 |
KIdsL1-1 |
0.69 |
0.78 |
0.44 |
KIdsL1-3 |
0.69 |
0.8 |
0.52 |
KIdsL1-5 |
0.69 |
0.74 |
0.48 |
sequences of five vowels resp. CV syllables which were displayed by DP were considered to provide the “reference”; output of the model yielded the hypotheses
data provided by L1 during three testing phases on days 1, 3 and 5 were evaluated by means of 5 different models
DeepSpeech_de = baseline model; KIds-0 : Deepspeech_de fine-tuned with kidsTALC; KIds-L1-1 : KIds-0 fine-tuned with data provided by L1 during day 1 learning phase;KIds-L1-3 : KIds-1 fine-tuned with data provided by L1 during day 3 learning phase;KIds-L1-5 : KIds-5 fine-tuned with data provided by L1 during day 5 learning phase
NOTE: WER-decrease between rows corresponds to increase of accuracy of the ASR model; WER-decrease between columns points to increase in L1's reading competence