System and Method for Adapting Automatic Speech Recognition Pronunciation by Acoustic Model Restructuring
First Claim
1. A method comprising:
- identifying an acoustic model, wherein the acoustic model is trained on native speech in a target dialect; and
replacing, via a processor, a phoneme in the acoustic model with a modified phoneme, wherein the modified phoneme is a weighted sum of plausible phonemes in a lattice of plausible phonemes associated with a class of a speaker.
2 Assignments
0 Petitions
Accused Products
Abstract
Disclosed herein are systems, computer-implemented methods, and computer-readable storage media for recognizing speech by adapting automatic speech recognition pronunciation by acoustic model restructuring. The method identifies an acoustic model and a matching pronouncing dictionary trained on typical native speech in a target dialect. The method collects speech from a new speaker resulting in collected speech and transcribes the collected speech to generate a lattice of plausible phonemes. Then the method creates a custom speech model for representing each phoneme used in the pronouncing dictionary by a weighted sum of acoustic models for all the plausible phonemes, wherein the pronouncing dictionary does not change, but the model of the acoustic space for each phoneme in the dictionary becomes a weighted sum of the acoustic models of phonemes of the typical native speech. Finally the method includes recognizing via a processor additional speech from the target speaker using the custom speech model.
-
Citations
20 Claims
-
1. A method comprising:
-
identifying an acoustic model, wherein the acoustic model is trained on native speech in a target dialect; and replacing, via a processor, a phoneme in the acoustic model with a modified phoneme, wherein the modified phoneme is a weighted sum of plausible phonemes in a lattice of plausible phonemes associated with a class of a speaker. - View Dependent Claims (2, 3, 4, 5, 6, 7)
-
-
8. A system comprising:
-
a processor; and a computer-readable storage medium having instructions stored which, when executed by the processor, result in the processor performing operations comprising; identifying an acoustic model, wherein the acoustic model is trained on native speech in a target dialect; and replacing a phoneme in the acoustic model with a modified phoneme, wherein the modified phoneme is a weighted sum of plausible phonemes in a lattice of plausible phonemes associated with a class of a speaker. - View Dependent Claims (9, 10, 11, 12, 13, 14)
-
-
15. A computer-readable storage device having instructions stored which, when executed by a computing device, result in the computing device performing operations comprising:
-
identifying an acoustic model, wherein the acoustic model is trained on native speech in a target dialect; and replacing a phoneme in the acoustic model with a modified phoneme, wherein the modified phoneme is a weighted sum of plausible phonemes in a lattice of plausible phonemes associated with a class of a speaker. - View Dependent Claims (16, 17, 18, 19, 20)
-
Specification