Artificial intelligence-based text-to-speech system and method

US 10,373,605 B2
Filed: 06/29/2018
Issued: 08/06/2019
Est. Priority Date: 05/18/2017
Status: Active Grant

First Claim

Patent Images

1. A text-to-speech (TTS) training system comprising:

a subsystem configured to receive an input vector from conversion of text, the subsystem including a neural network interacting with a pre-existing knowledgebase of phonemes to apply an error signal to correct for speech signal distortions of the pre-existing knowledgebase of phonemes to generate a corrected speech signal; and

a training subsystem coupled to the subsystem, the training subsystem configured to iteratively correct the subsystem for the speech signal distortions of the pre-existing knowledgebase of phonemes based on psychoacoustic processing for the subsystem to apply the error signal to correct for the speech signal distortions of the pre-existing knowledgebase of phonemes to generate the corrected speech signal, wherein the training subsystem is further configured to ignore inaudible errors of the corrected speech signal based on masking.

View all claims

3 Assignments

Timeline View

Assignment View

0 Petitions

Accused Products

Abstract

A technique improves training and speech quality of a text-to-speech (TTS) system having an artificial intelligence, such as a neural network. The TTS system is organized as a front-end subsystem and a back-end subsystem. The front-end subsystem is configured to provide analysis and conversion of text into input vectors, each having at least a base frequency, f₀, a phenome duration, and a phoneme sequence that is processed by a signal generation unit of the back-end subsystem. The signal generation unit includes the neural network interacting with a pre-existing knowledgebase of phenomes to generate audible speech from the input vectors. The technique applies an error signal from the neural network to correct imperfections of the pre-existing knowledgebase of phenomes to generate audible speech signals. A back-end training system is configured to train the signal generation unit by applying psychoacoustic principles to improve quality of the generated audible speech signals.

20 Citations

View as Search Results

20 Claims

1. A text-to-speech (TTS) training system comprising:
- a subsystem configured to receive an input vector from conversion of text, the subsystem including a neural network interacting with a pre-existing knowledgebase of phonemes to apply an error signal to correct for speech signal distortions of the pre-existing knowledgebase of phonemes to generate a corrected speech signal; and
  
  a training subsystem coupled to the subsystem, the training subsystem configured to iteratively correct the subsystem for the speech signal distortions of the pre-existing knowledgebase of phonemes based on psychoacoustic processing for the subsystem to apply the error signal to correct for the speech signal distortions of the pre-existing knowledgebase of phonemes to generate the corrected speech signal, wherein the training subsystem is further configured to ignore inaudible errors of the corrected speech signal based on masking.
- View Dependent Claims (2, 3, 4, 5, 6, 7, 8, 9, 10)
- - 2. The TTS training system of claim 1 wherein the training subsystem is further configured to calculate audible errors of the corrected speech signal by the psychoacoustic processing.
  - 3. The TTS training system of claim 1 wherein masking includes one of frequency masking and temporal masking.
  - 4. The TTS training system of claim 1 wherein the training subsystem is further configured to iteratively modify the neural network to correct the subsystem.
  - 5. The TTS training system of claim 1 further comprising a psychoacoustic generator configured to analyze a reference audio signal to determine masking information.
  - 6. The TTS training system of claim 5 wherein the psychoacoustic generator is further configured to identify locations and energy levels that are audible and inaudible.
  - 7. The TTS training system of claim 1 further comprising a quality indicator calculator configured to determine a quality indicator based on audible errors using the psychoacoustic processing to ignore inaudible errors.
  - 8. The TTS training system of claim 7 wherein the quality indicator is calculated based on a total of audible error signal energy.
  - 9. The TTS training system of claim 7 wherein the iterative correction further comprises a modification of the neural network that is performed so that a total audible error signal energy is below a quality threshold.
  - 10. The TTS training system of claim 9 wherein the quality indicator at least converges close to zero with the modification of the neural network.

11. A method of training text-to-speech (TTS) processing comprising:
- receiving, by a subsystem, an input vector from conversion of text;
  
  interacting, by a neural network of the subsystem, with a pre-existing knowledgebase of phonemes to apply an error signal to correct for speech signal distortions of the pre-existing knowledgebase of phonemes to generate a corrected speech signal;
  
  iteratively correcting, by a training subsystem coupled to the subsystem, the subsystem for the speech signal distortions of the pre-existing knowledgebase of phonemes based on psychoacoustic processing for the subsystem to apply the error signal to correct for the speech signal distortions of the pre-existing knowledgebase of phonemes to generate the corrected speech signal; and
  
  ignoring, by the training subsystem, inaudible errors of the corrected speech signal based on masking.
- View Dependent Claims (12, 13, 14, 15, 16, 17, 18, 19)
- - 12. The method of training TTS processing of claim 11 further comprising calculating audible errors of the corrected speech signal by the psychoacoustic processing.
  - 13. The method of training TTS processing of claim 11 wherein masking includes one of frequency masking and temporal masking.
  - 14. The method of training TTS processing of claim 11 wherein the iterative correction further comprises modifying the neural network to correct the subsystem.
  - 15. The method of training TTS processing of claim 11 further comprising analyzing a reference audio signal to determine masking information.
  - 16. The method of training TTS processing of claim 11 further comprising identifying locations and energy levels that are audible and inaudible.
  - 17. The method of training TTS processing of claim 11 further comprising:
    - determining a quality indicator based on audible errors; and
      
      ignoring inaudible errors using the psychoacoustic processing.
  - 18. The method of training TTS processing of claim 17 wherein the quality indicator is calculated based on a total of audible error signal energy.
  - 19. The method of training TTS processing of claim 17 wherein the iterative correction further comprises modifying the neural network so that a total audible error signal energy is below a quality threshold.

20. A non-transitory computer-readable medium having program instructions for training text-to-speech (TTS) processing which, when executed across one or more processors, causes at least a portion of the one or more processors to perform operations comprising:
- receiving, by a subsystem, an input vector from conversion of text;
  
  interacting, by a neural network of the subsystem, with a pre-existing knowledgebase of phonemes to apply an error signal to correct for speech signal distortions of the pre-existing knowledgebase of phonemes to generate a corrected speech signal;
  
  iteratively correcting, by a training subsystem coupled to the subsystem, the subsystem for the speech signal distortions of the pre-existing knowledgebase of phonemes based on psychoacoustic processing to apply the error signal to correct for the speech signal distortions of the pre-existing knowledgebase of phonemes to generate the corrected speech signal; and
  
  ignoring, by the training subsystem, inaudible errors of the corrected speech signal based on masking.

Specification

Resources

Litigation Campaign Assessment

Current Assignee
Telepathy Labs, Inc.
Original Assignee
Telepathy Labs, Inc.
Inventors
Reber, Martin, Avijeet, Vijeta
Primary Examiner(s)
Jackson, Jakieda R

Application Number

US16/022,823
Publication Number

US 20180336882A1
Time in Patent Office

403 Days
Field of Search

7042001, 704232, 704259
US Class Current
CPC Class Codes

G06F 18/10   Pre-processing; Data cleansing

G06F 18/2135   based on approximation crit...

G06F 18/217   Validation; Performance eva...

G06N 3/02   Neural networks

G06N 3/042   Knowledge-based neural netw...

G06N 3/08   Learning methods

G06N 5/02   Knowledge representation; S...

G10L 13/04   Details of speech synthesis...

G10L 13/08   Text analysis or generation...

G10L 19/00   Speech or audio signals ana...

Artificial intelligence-based text-to-speech system and method

First Claim

3 Assignments

0 Petitions

Accused Products

Abstract

20 Citations

20 Claims

Specification

Use Cases

Quick Links

Others

Artificial intelligence-based text-to-speech system and method

First Claim

3 Assignments

Subscription Required

Subscription Required

0 Petitions

Subscription Required

Accused Products

Subscription Required

Abstract

20 Citations

20 Claims

Specification

Subscription Required

Use Cases

Quick Links

Others