Method and system for voice recognition employing multiple voice-recognition techniques

US 9,570,076 B2
Filed: 02/22/2013
Issued: 02/14/2017
Est. Priority Date: 10/30/2012
Status: Active Grant

First Claim

Patent Images

1. A computer-implemented method comprising:

receiving audio data that encodes an utterance;

obtaining, as a result of performing speech-to-text voice recognition on the audio data, a first transcription of the utterance;

segmenting the first transcription into two or more discrete terms;

determining that a first particular term from among the two or more discrete terms is included among a predefined set of terms that are associated with a word spotting process that involves determining whether an acoustic fingerprint of a given portion of audio data is an acoustic match with one or more given terms without performing speech-to-text voice recognition;

determining that the two or more discrete terms other than the first particular term are included among an additional predefined set of terms that are associated with the predefined set of terms that are associated with a word spotting process;

in response to determining that the two or more discrete terms other than the first particular term are included among the additional predefined set of terms that are associated with the predefined set of terms that are associated with the word spotting process, obtaining, as a result of performing the word spotting process on a portion of the audio data that corresponds to a second particular term from among the two or more discrete terms other than the first particular term without re-performing speech-to-text voice recognition on the portion of the audio data, an indication that an acoustic fingerprint associated with the portion of the audio data that corresponds to the second particular term is an acoustic match with one or more terms of the predefined set of terms that are associated with the word spotting process;

obtaining, as a result of re-performing speech-to-text voice recognition on a portion of the audio data that does not correspond to the second particular term, a second transcription of the utterance using the portion of the audio data that does not correspond to the second particular term;

generating a third transcription of the utterance based at least on (i) the second transcription of the utterance that was obtained as a result of re-performing speech-to-text voice recognition on the portion of the audio data that does not correspond to the second particular term, and (ii) the one or more terms of the predefined set of terms that are indicated, as a result of performing the word spotting process on the portion of the audio data that corresponds to the second particular term without re-performing speech-to-text voice recognition of the audio data, as an acoustic match with the portion of the audio data that corresponds to the second particular term; and

providing the third transcription of the utterance for output.

View all claims

2 Assignments

Timeline View

Assignment View

0 Petitions

Accused Products

Abstract

A method and system for voice recognition are disclosed. In one example embodiment, the method includes receiving voice input information by way of a receiver on a mobile device and performing, by way of at least one processing device on the mobile device, first and second processing operations respectively with respect to first and second voice input portions, respectively, which respectively correspond to and are based at least indirectly upon different respective portions of the voice input information. The first processing operation includes a speech-to-text operation and the second processing operation includes an alternate processing operation. Additionally, the method includes generating recognized voice information based at least indirectly upon results from the first and second processing operations, and performing at least one action based at least in part upon the recognized voice information, where the at least one action includes outputting at least one signal by an output device.

61 Citations

View as Search Results

20 Claims

1. A computer-implemented method comprising:
- receiving audio data that encodes an utterance;
  
  obtaining, as a result of performing speech-to-text voice recognition on the audio data, a first transcription of the utterance;
  
  segmenting the first transcription into two or more discrete terms;
  
  determining that a first particular term from among the two or more discrete terms is included among a predefined set of terms that are associated with a word spotting process that involves determining whether an acoustic fingerprint of a given portion of audio data is an acoustic match with one or more given terms without performing speech-to-text voice recognition;
  
  determining that the two or more discrete terms other than the first particular term are included among an additional predefined set of terms that are associated with the predefined set of terms that are associated with a word spotting process;
  
  in response to determining that the two or more discrete terms other than the first particular term are included among the additional predefined set of terms that are associated with the predefined set of terms that are associated with the word spotting process, obtaining, as a result of performing the word spotting process on a portion of the audio data that corresponds to a second particular term from among the two or more discrete terms other than the first particular term without re-performing speech-to-text voice recognition on the portion of the audio data, an indication that an acoustic fingerprint associated with the portion of the audio data that corresponds to the second particular term is an acoustic match with one or more terms of the predefined set of terms that are associated with the word spotting process;
  
  obtaining, as a result of re-performing speech-to-text voice recognition on a portion of the audio data that does not correspond to the second particular term, a second transcription of the utterance using the portion of the audio data that does not correspond to the second particular term;
  
  generating a third transcription of the utterance based at least on (i) the second transcription of the utterance that was obtained as a result of re-performing speech-to-text voice recognition on the portion of the audio data that does not correspond to the second particular term, and (ii) the one or more terms of the predefined set of terms that are indicated, as a result of performing the word spotting process on the portion of the audio data that corresponds to the second particular term without re-performing speech-to-text voice recognition of the audio data, as an acoustic match with the portion of the audio data that corresponds to the second particular term; and
  
  providing the third transcription of the utterance for output.
- View Dependent Claims (2, 3, 4, 5, 6, 7)
- - 2. The method of claim 1, wherein segmenting the first transcription into two or more discrete terms is based on a grammar structure of the first transcription.
  - 3. The method of claim 1, wherein the predefined set of terms that are associated with the word spotting process includes terms entered by a user.
  - 4. The method of claim 1, wherein the first transcription and the second transcription are obtained using different speech-to-text algorithms.
  - 5. The method of claim 1, wherein providing the third transcription of the utterance for output comprises:
    - providing the third transcription of the utterance to an application.
  - 6. The method of claim 1, wherein generating the third transcription comprises:
    - concatenating (i) the second transcription of the utterance that was obtained as a result of re-performing speech-to-text voice recognition on the portion of the audio data that does not correspond to the second particular term, and (ii) the one or more terms of the predefined set of terms that are indicated, as a result of performing the word spotting process on the portion of the audio data that corresponds to the second particular term without re-performing speech-to-text voice recognition of the audio data, as an acoustic match with the portion of the audio data that corresponds to the second particular term.
  - 7. The method of claim 1, comprising:
    - performing the word spotting process on a portion of the audio data that corresponds to the second particular term.

8. A system comprising:
- one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising;
  
  receiving audio data that encodes an utterance;
  
  obtaining, as a result of performing speech-to-text voice recognition on the audio data, a first transcription of the utterance;
  
  segmenting the first transcription into two or more discrete terms;
  
  determining that a first particular term from among the two or more discrete terms is included among a predefined set of terms that are associated with a word spotting process that involves determining whether an acoustic fingerprint of a given portion of audio data is an acoustic match with one or more given terms without performing speech-to-text voice recognition;
  
  determining that the two or more discrete terms other than the first particular term are included among an additional predefined set of terms that are associated with the predefined set of terms that are associated with a word spotting process;
  
  in response to determining that the two or more discrete terms other than the first particular term are included among the additional predefined set of terms that are associated with the predefined set of terms that are associated with the word spotting process, obtaining, as a result of performing the word spotting process on a portion of the audio data that corresponds to a second particular term from among the two or more discrete terms other than the first particular term without re-performing speech-to-text voice recognition on the portion of the audio data, an indication that an acoustic fingerprint associated with the portion of the audio data that corresponds to the second particular term is an acoustic match with one or more terms of the predefined set of terms that are associated with the word spotting process;
  
  obtaining, as a result of re-performing speech-to-text voice recognition on a portion of the audio data that does not correspond to the second particular term, a second transcription of the utterance using the portion of the audio data that does not correspond to the second particular term;
  
  generating a third transcription of the utterance based at least on (i) the second transcription of the utterance that was obtained as a result of re-performing speech-to-text voice recognition on the portion of the audio data that does not correspond to the second particular term, and (ii) the one or more terms of the predefined set of terms that are indicated, as a result of performing the word spotting process on the portion of the audio data that corresponds to the second particular term without re-performing speech-to-text voice recognition of the audio data, as an acoustic match with the portion of the audio data that corresponds to the second particular term; and
  
  providing the third transcription of the utterance for output.
- View Dependent Claims (9, 10, 11, 12, 13, 14)
- - 9. The system of claim 8, wherein segmenting the first transcription into two or more discrete terms is based on a grammar structure of the first transcription.
  - 10. The system of claim 8, wherein the predefined set of terms that are associated with the word spotting process includes terms entered by a user.
  - 11. The system of claim 8, wherein the first transcription and the second transcription are obtained using different speech-to-text algorithms.
  - 12. The system of claim 8, wherein providing the third transcription of the utterance for output comprises:
    - providing the third transcription of the utterance to an application.
  - 13. The system of claim 8, wherein generating the third transcription comprises:
    - concatenating (i) the second transcription of the utterance that was obtained as a result of re-performing speech-to-text voice recognition on the portion of the audio data that does not correspond to the second particular term, and (ii) the one or more terms of the predefined set of terms that are indicated, as a result of performing the word spotting process on the portion of the audio data that corresponds to the second particular term without re-performing speech-to-text voice recognition of the audio data, as an acoustic match with the portion of the audio data that corresponds to the second particular term.
  - 14. The system of claim 8, wherein the operations further comprise:
    - performing the word spotting process on a portion of the audio data that corresponds to the second particular term.

15. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:
- receiving audio data that encodes an utterance;
  
  obtaining, as a result of performing speech-to-text voice recognition on the audio data, a first transcription of the utterance;
  
  segmenting the first transcription into two or more discrete terms;
  
  determining that a first particular term from among the two or more discrete terms is included among a predefined set of terms that are associated with a word spotting process that involves determining whether an acoustic fingerprint of a given portion of audio data is an acoustic match with one or more given terms without performing speech-to-text voice recognition;
  
  determining that the two or more discrete terms other than the first particular term are included among an additional predefined set of terms that are associated with the predefined set of terms that are associated with a word spotting process;
  
  in response to determining that the two or more discrete terms other than the first particular term are included among the additional predefined set of terms that are associated with the predefined set of terms that are associated with the word spotting process, obtaining, as a result of performing the word spotting process on a portion of the audio data that corresponds to a second particular term from among the two or more discrete terms other than the first particular term without re-performing speech-to-text voice recognition on the portion of the audio data, an indication that an acoustic fingerprint associated with the portion of the audio data that corresponds to the second particular term is an acoustic match with one or more terms of the predefined set of terms that are associated with the word spotting process;
  
  obtaining, as a result of re-performing speech-to-text voice recognition on a portion of the audio data that does not correspond to the second particular term, a second transcription of the utterance using the portion of the audio data that does not correspond to the second particular term;
  
  generating a third transcription of the utterance based at least on (i) the second transcription of the utterance that was obtained as a result of re-performing speech-to-text voice recognition on the portion of the audio data that does not correspond to the second particular term, and (ii) the one or more terms of the predefined set of terms that are indicated, as a result of performing the word spotting process on the portion of the audio data that corresponds to the second particular term without re-performing speech-to-text voice recognition of the audio data, as an acoustic match with the portion of the audio data that corresponds to the second particular term; and
  
  providing the third transcription of the utterance for output.
- View Dependent Claims (16, 17, 18, 19, 20)
- - 16. The medium of claim 15, wherein segmenting the first transcription into two or more discrete terms is based on a grammar structure of the first transcription.
  - 17. The medium of claim 15, wherein the first transcription and the second transcription are obtained using different speech-to-text algorithms.
  - 18. The medium of claim 15, wherein providing the third transcription of the utterance for output comprises:
    - providing the third transcription of the utterance to an application.
  - 19. The medium of claim 15, wherein generating the third transcription comprises:
    - concatenating (i) the second transcription of the utterance that was obtained as a result of re-performing speech-to-text voice recognition on the portion of the audio data that does not correspond to the second particular term, and (ii) the one or more terms of the predefined set of terms that are indicated, as a result of performing the word spotting process on the portion of the audio data that corresponds to the second particular term without re-performing speech-to-text voice recognition of the audio data, as an acoustic match with the portion of the audio data that corresponds to the second particular term.
  - 20. The medium of claim 15, wherein the operations further comprise:
    - performing the word spotting process on a portion of the audio data that corresponds to the second particular term.

Specification

Resources

Litigation Campaign Assessment

Current Assignee
Google Technology Holdings LLC (Alphabet Inc.)
Original Assignee
Google Technology Holdings LLC (Alphabet Inc.)
Inventors
Sierawski, Jeffrey A, Labowicz, Michael P, Bekkerman, Boris, Ondo, Kazuhiro
Primary Examiner(s)
Goddard, Tammy Paige
Assistant Examiner(s)
Leland, III, Edwin S

Application Number

US13/774,398
Publication Number

US 20140122071A1
Time in Patent Office

1,453 Days
Field of Search

704/235
US Class Current

1/1
CPC Class Codes

G06F 3/167   Audio in a user interface, ...

G10L 15/18   using natural language mode...

G10L 15/26   Speech to text systems G10L...

G10L 15/32   Multiple recognisers used i...

G10L 2015/088   Word spotting

H04M 1/724   User interfaces specially a...

H04M 1/72436   for text messaging, e.g. sh...

H04M 2250/74   with voice recognition mean...

Method and system for voice recognition employing multiple voice-recognition techniques

First Claim

2 Assignments

0 Petitions

Accused Products

Abstract

61 Citations

20 Claims

Specification

Solutions

Use Cases

Quick Links

Method and system for voice recognition employing multiple voice-recognition techniques

First Claim

2 Assignments

Subscription Required

Subscription Required

0 Petitions

Subscription Required

Accused Products

Subscription Required

Abstract

61 Citations

20 Claims

Specification

Subscription Required

Solutions

Use Cases

Quick Links