System and method for synthetically generated speech describing media content

US 9,558,735 B2
Filed: 03/18/2016
Issued: 01/31/2017
Est. Priority Date: 06/06/2008
Status: Active Grant

First Claim

Patent Images

1. A method comprising:

accessing, by a system including a processor, user input that was captured during a presentation of media content, wherein the user input comprises a metadata request associated with the media content;

obtaining, by the system according to the metadata request, metadata for output; and

outputting, by the system, synthetically generated speech associated with the media content during an audio gap in the presentation of the media content, the synthetically generated speech having an accent selected from a plurality of accents based on the metadata.

View all claims

2 Assignments

Timeline View

Assignment View

0 Petitions

Accused Products

Abstract

Disclosed herein are systems, methods, and computer readable-media for providing an automatic synthetically generated voice describing media content, the method comprising receiving one or more pieces of metadata for a primary media content, selecting at least one piece of metadata for output, and outputting the at least one piece of metadata as synthetically generated speech with the primary media content. Other aspects of the invention involve alternative output, output speech simultaneously with the primary media content, output speech during gaps in the primary media content, translate metadata in foreign language, tailor voice, accent, and language to match the metadata and/or primary media content. A user may control output via a user interface or output may be customized based on preferences in a user profile.

Citations

20 Claims

1. A method comprising:
- accessing, by a system including a processor, user input that was captured during a presentation of media content, wherein the user input comprises a metadata request associated with the media content;
  
  obtaining, by the system according to the metadata request, metadata for output; and
  
  outputting, by the system, synthetically generated speech associated with the media content during an audio gap in the presentation of the media content, the synthetically generated speech having an accent selected from a plurality of accents based on the metadata.
- View Dependent Claims (2, 3, 4, 5, 6, 7, 8, 9, 10, 11)
- - 2. The method of claim 1, wherein the synthetically generated speech has a language, a lexicon, or a combination thereof selected from a plurality of languages, lexicons or combinations thereof according to a user profile.
  - 3. The method of claim 1, wherein the user input comprises a gesture.
  - 4. The method of claim 1, wherein the user input comprises an oral command.
  - 5. The method of claim 1, wherein the outputting of the synthetically generated speech is based on a user profile.
  - 6. The method of claim 1, wherein the media content comprises a group of songs, and wherein the audio gap is in-between songs of the group of songs.
  - 7. The method of claim 1, further comprising analyzing the media content to determine tone and prosody which indicate the accent.
  - 8. The method of claim 1, wherein the obtaining of the metadata is based on a selection from among a group of metadata received by the system with the media content.
  - 9. The method of claim 1, wherein the obtaining of the metadata is based on a selection from among a group of metadata stored by the system.
  - 10. The method of claim 1, wherein the obtaining of the metadata is based on a selection from among a group of metadata accessed by the system from a remote metadata source.
  - 11. The method of claim 1, further comprising:
    - determining the metadata is in a foreign language, wherein the accent corresponds to the foreign language; and
      
      translating the metadata to another language from the foreign language before output.

12. A system comprising:
- a processor; and
  
  a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising;
  
  receiving a gesture from a user during a presentation of media content, wherein the gesture comprises a metadata request associated with the media content;
  
  accessing a group of metadata from a source that is different from the source of the media content;
  
  selecting, according to the metadata request, metadata from the group of metadata; and
  
  outputting synthetically generated speech according to the metadata during an audio gap in the presentation of the media content, the synthetically generated speech having an accent, a language, a lexicon, or a combination thereof selected from a plurality of accents, languages, lexicons or combinations thereof.
- View Dependent Claims (13, 14, 15, 16)
- - 13. The system of claim 12, wherein the gesture is accompanied by an oral command.
  - 14. The system of claim 12, wherein the outputting of the synthetically generated speech is based on a user profile.
  - 15. The system of claim 12, wherein the source of the metadata is a local storage accessible to the processor.
  - 16. The system of claim 12, wherein the source of the metadata is a remote storage accessible to the processor.

17. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:
- receiving a gesture from a user during a presentation of media content, wherein the gesture comprises a metadata request associated with the media content;
  
  selecting, according to the metadata request, metadata from a group of metadata;
  
  accessing a user profile for the user; and
  
  outputting synthetically generated speech according to the metadata and the user profile during an audio gap in the presentation of the media content, the synthetically generated speech having an accent, a language, a lexicon, or a combination thereof selected from a plurality of accents, languages, lexicons or combinations thereof.
- View Dependent Claims (18, 19, 20)
- - 18. The computer-readable storage device of claim 17, wherein the operations further comprise accessing the group of metadata from a source that is different from the source of the media content.
  - 19. The computer-readable storage device of claim 18, wherein the source of the metadata is a remote storage accessible to the computing device.
  - 20. The computer-readable storage device of claim 18, wherein the source of the metadata is a local storage accessible to the computing device.

Specification

Resources

Litigation Campaign Assessment

Current Assignee
AT&T Intellectual Property I LP (AT&T, Inc.)
Original Assignee
AT&T Intellectual Property I LP (AT&T, Inc.)
Inventors
Schroeter, Horst J, Roberts, Linda, Nguyen, Hong Thi
Primary Examiner(s)
PULLIAS, JESSE SCOTT

Application Number

US15/074,362
Publication Number

US 20160203815A1
Time in Patent Office

319 Days
Field of Search

704231-275
US Class Current

1/1
CPC Class Codes

G06F 3/017   Gesture based interaction, ...

G06F 3/04842   Selection of displayed obje...

G06F 3/167   Audio in a user interface, ...

G06F 40/58   Use of machine translation,...

G10L 13/00   Speech synthesis; Text to s...

G10L 13/033   Voice editing, e.g. manipul...

G10L 13/04   Details of speech synthesis...

G10L 13/086   Detection of language

G10L 13/10   Prosody rules derived from ...

G10L 15/22   Procedures used during a sp...

G10L 2013/083   Special characters, e.g. pu...

System and method for synthetically generated speech describing media content

First Claim

2 Assignments

0 Petitions

Accused Products

Abstract

Citations

20 Claims

Specification

Solutions

Use Cases

Quick Links

System and method for synthetically generated speech describing media content

First Claim

2 Assignments

Subscription Required

Subscription Required

0 Petitions

Subscription Required

Accused Products

Subscription Required

Abstract

Citations

20 Claims

Specification

Subscription Required

Solutions

Use Cases

Quick Links