End of query detection

US 10,593,352 B2
Filed: 06/06/2018
Issued: 03/17/2020
Est. Priority Date: 06/06/2017
Status: Active Grant

First Claim

Patent Images

1. A computer-implemented method comprising:

receiving audio data that corresponds to an utterance spoken by a user;

applying, to the audio data, an end of query model that (i) is configured to determine a confidence score that reflects a likelihood that the utterance is a complete utterance and (ii) was trained using audio data from complete utterances and from incomplete utterances;

based on applying the end of query model that (i) is configured to determine the confidence score that reflects the likelihood that the utterance is a complete utterance and (ii) was trained using the audio data from the complete utterances and from the incomplete utterances, determining the confidence score that reflects a likelihood that the utterance is a complete utterance;

comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to a confidence score threshold;

based on comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold, determining whether the utterance is likely complete or likely incomplete; and

based on determining whether the utterance is likely complete or likely incomplete, providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance.

View all claims

2 Assignments

Timeline View

Assignment View

0 Petitions

Accused Products

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for detecting an end of a query are disclosed. In one aspect, a method includes the actions of receiving audio data that corresponds to an utterance spoken by a user. The actions further include applying, to the audio data, an end of query model. The actions further include determining the confidence score that reflects a likelihood that the utterance is a complete utterance. The actions further include comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to a confidence score threshold. The actions further include determining whether the utterance is likely complete or likely incomplete. The actions further include providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance.

85 Citations

View as Search Results

20 Claims

1. A computer-implemented method comprising:
- receiving audio data that corresponds to an utterance spoken by a user;
  
  applying, to the audio data, an end of query model that (i) is configured to determine a confidence score that reflects a likelihood that the utterance is a complete utterance and (ii) was trained using audio data from complete utterances and from incomplete utterances;
  
  based on applying the end of query model that (i) is configured to determine the confidence score that reflects the likelihood that the utterance is a complete utterance and (ii) was trained using the audio data from the complete utterances and from the incomplete utterances, determining the confidence score that reflects a likelihood that the utterance is a complete utterance;
  
  comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to a confidence score threshold;
  
  based on comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold, determining whether the utterance is likely complete or likely incomplete; and
  
  based on determining whether the utterance is likely complete or likely incomplete, providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance.
- View Dependent Claims (2, 3, 4, 5, 6, 7, 8, 9)
- - 2. The method of claim 1, comprising:
    - based on comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold, determining that the confidence score satisfies the confidence score threshold,wherein determining whether the utterance is likely complete or likely incomplete comprises determining the utterance is likely complete based on determining that the confidence score satisfies the confidence score threshold,wherein providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance comprises providing, for output, the instruction to deactivate the microphone that is receiving the utterance,generating a transcription of the audio data, andproviding, for output, the transcription.
  - 3. The method of claim 2, comprising:
    - receiving, from a user, data confirming that the user finished speaking; and
      
      based on receiving the data confirming that the user finished speaking, updating the end of query model.
  - 4. The method of claim 1, comprising:
    - based on comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold, determining that the confidence score does not satisfy the confidence score threshold,wherein determining whether the utterance is likely complete or likely incomplete comprises determining the utterance is likely incomplete based on determining that the confidence score does not satisfy the confidence score threshold, andwherein providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance comprises providing, for output, the instruction to maintain the microphone in an active state.
  - 5. The method of claim 1, comprising:
    - receiving audio data of multiple complete utterances and multiple incomplete utterances; and
      
      training, using machine learning, the end of query model using the audio data of the multiple complete utterances and the multiple incomplete utterances.
  - 6. The method of claim 1, wherein the end of query model is configured to determine the confidence score that reflects the likelihood that the utterance is a complete utterance based on acoustic speech characteristics of the utterance that include pitch, loudness, intonation, sharpness, articulation, roughness, instability, and speech rate.
  - 7. The method of claim 1, comprising:
    - determining that a speech decoder that is configured to generate a transcription of the audio data and that is configured to determine whether the utterance is likely complete or likely incomplete has not determined whether the utterance is likely complete or likely incomplete,wherein determining whether the utterance is likely complete or likely incomplete is based on only comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold.
  - 8. The method of claim 7, wherein the speech decoder uses a language model to determine whether the utterance is likely complete or likely incomplete.
  - 9. The method of claim 1, comprising:
    - determining that a speech decoder that is configured to generate a transcription of the audio data and that is configured to determine whether the utterance is likely complete or likely incomplete has determined whether the utterance is likely complete or likely incomplete,wherein determining whether the utterance is likely complete or likely incomplete is based on (i) comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold and (ii) the speech decoder determining whether the utterance is likely complete or likely incomplete.

10. A system comprising:
- one or more computers; and
  
  one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising;
  
  receiving audio data that corresponds to an utterance spoken by a user;
  
  applying, to the audio data, an end of query model that (i) is configured to determine a confidence score that reflects a likelihood that the utterance is a complete utterance and (ii) was trained using audio data from complete utterances and from incomplete utterances;
  
  based on applying the end of query model that (i) is configured to determine the confidence score that reflects the likelihood that the utterance is a complete utterance and (ii) was trained using the audio data from the complete utterances and from the incomplete utterances, determining the confidence score that reflects a likelihood that the utterance is a complete utterance;
  
  comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to a confidence score threshold;
  
  based on comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold, determining whether the utterance is likely complete or likely incomplete; and
  
  based on determining whether the utterance is likely complete or likely incomplete, providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance.
- View Dependent Claims (11, 12, 13, 14, 15, 16, 17, 18)
- - 11. The system of claim 10, wherein the operations further comprise:
    - based on comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold, determining that the confidence score satisfies the confidence score threshold,wherein determining whether the utterance is likely complete or likely incomplete comprises determining the utterance is likely complete based on determining that the confidence score satisfies the confidence score threshold,wherein providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance comprises providing, for output, the instruction to deactivate the microphone that is receiving the utterance,generating a transcription of the audio data, andproviding, for output, the transcription.
  - 12. The system of claim 11, wherein the operations further comprise:
    - receiving, from a user, data confirming that the user finished speaking; and
      
      based on receiving the data confirming that the user finished speaking, updating the end of query model.
  - 13. The system of claim 10, wherein the operations further comprise:
    - based on comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold, determining that the confidence score does not satisfy the confidence score threshold,wherein determining whether the utterance is likely complete or likely incomplete comprises determining the utterance is likely incomplete based on determining that the confidence score does not satisfy the confidence score threshold, andwherein providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance comprises providing, for output, the instruction to maintain the microphone in an active state.
  - 14. The system of claim 10, wherein the operations further comprise:
    - receiving audio data of multiple complete utterances and multiple incomplete utterances; and
      
      training, using machine learning, the end of query model using the audio data of the multiple complete utterances and the multiple incomplete utterances.
  - 15. The system of claim 10, wherein the end of query model is configured to determine the confidence score that reflects the likelihood that the utterance is a complete utterance based on acoustic speech characteristics of the utterance that include pitch, loudness, intonation, sharpness, articulation, roughness, instability, and speech rate.
  - 16. The system of claim 10, wherein the operations further comprise:
    - determining that a speech decoder that is configured to generate a transcription of the audio data and that is configured to determine whether the utterance is likely complete or likely incomplete has not determined whether the utterance is likely complete or likely incomplete,wherein determining whether the utterance is likely complete or likely incomplete is based on only comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold.
  - 17. The system of claim 16, wherein the speech decoder uses a language model to determine whether the utterance is likely complete or likely incomplete.
  - 18. The system of claim 10, wherein the operations further comprise:
    - determining that a speech decoder that is configured to generate a transcription of the audio data and that is configured to determine whether the utterance is likely complete or likely incomplete has determined whether the utterance is likely complete or likely incomplete,wherein determining whether the utterance is likely complete or likely incomplete is based on (i) comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold and (ii) the speech decoder determining whether the utterance is likely complete or likely incomplete.

19. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:
- receiving audio data that corresponds to an utterance spoken by a user;
  
  applying, to the audio data, an end of query model that (i) is configured to determine a confidence score that reflects a likelihood that the utterance is a complete utterance and (ii) was trained using audio data from complete utterances and from incomplete utterances;
  
  based on applying the end of query model that (i) is configured to determine the confidence score that reflects the likelihood that the utterance is a complete utterance and (ii) was trained using the audio data from the complete utterances and from the incomplete utterances, determining the confidence score that reflects a likelihood that the utterance is a complete utterance;
  
  comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to a confidence score threshold;
  
  based on comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold, determining whether the utterance is likely complete or likely incomplete; and
  
  based on determining whether the utterance is likely complete or likely incomplete, providing, for output, an instruction to (i) maintain a microphone that is receiving the utterance in an active state or (ii) deactivate the microphone that is receiving the utterance.
- View Dependent Claims (20)
- - 20. The medium of claim 19, wherein the operations further comprise:
    - determining that a speech decoder that is configured to generate a transcription of the audio data and that is configured to determine whether the utterance is likely complete or likely incomplete has not determined whether the utterance is likely complete or likely incomplete,wherein determining whether the utterance is likely complete or likely incomplete is based on only comparing the confidence score that reflects the likelihood that the utterance is a complete utterance to the confidence score threshold.

Specification

Resources

Litigation Campaign Assessment

Current Assignee
Google LLC (Alphabet Inc.)
Original Assignee
Google LLC (Alphabet Inc.)
Inventors
Simko, Gabor, Parada San Martin, Maria Carolina, Shannon, Sean Matthew
Primary Examiner(s)
Vo, Huyen X

Application Number

US16/001,140
Publication Number

US 20180350395A1
Time in Patent Office

650 Days
Field of Search
US Class Current
CPC Class Codes

G10L 15/065   Adaptation

G10L 15/18   using natural language mode...

G10L 15/187   Phonemic context, e.g. pron...

G10L 15/22   Procedures used during a sp...

G10L 2025/783   based on threshold decision

G10L 25/78   Detection of presence or ab...

End of query detection

First Claim

2 Assignments

0 Petitions

Accused Products

Abstract

85 Citations

20 Claims

Specification

Use Cases

Quick Links

Others

End of query detection

First Claim

2 Assignments

Subscription Required

Subscription Required

0 Petitions

Subscription Required

Accused Products

Subscription Required

Abstract

85 Citations

20 Claims

Specification

Subscription Required

Use Cases

Quick Links

Others