Method and apparatus for processing video sequences

US 9,214,030 B2
Filed: 04/25/2008
Issued: 12/15/2015
Est. Priority Date: 05/07/2007
Status: Active Grant

First Claim

Patent Images

1. A method for processing a video sequence comprised of a plurality of frames, said method comprising:

extracting a feature from each of said frames;

determining correspondences between said extracted feature from said frames;

determining motion in said video sequence based on said determined correspondences, said determining motion using a modified random sample consensus algorithm that selects samples from buckets, iterates a model estimation multiple times including all inliers obtained so far in each iteration, and finds a model that maximizes a data likelihood, wherein a motion hypothesis is derived with a least squares method when the obtained inliers are determined to be less than a number, and the motion hypothesis is derived with a weighted total least squares method otherwise;

generating a forward warping matrix and a background warping matrix for each of said frames based on said determined motion;

generating a forward warping error and a backward warping error for each of said frames based on said forward warping matrix and said background warping matrix;

generating a foreground/background mask for each of said frames based on said forward warping error and said backward warping error; and

generating a background mosaic by mapping said frames to a common coordinate system; and

extracting foreground information from each of said frames based on said background mosaic.

View all claims

4 Assignments

Timeline View

Assignment View

0 Petitions

Accused Products

Abstract

A method for processing a video sequence having a plurality of frames includes the steps of: extracting features from each of the frames, determining correspondences between the extracted features from two of the frames, estimating motion in the video sequence based on the determined correspondences, generating a background mosaic for the video sequence based on the estimated motion, and performing foreground-background segmentation on each of the frames based on the background mosaic.

33 Citations

View as Search Results

18 Claims

1. A method for processing a video sequence comprised of a plurality of frames, said method comprising:
- extracting a feature from each of said frames;
  
  determining correspondences between said extracted feature from said frames;
  
  determining motion in said video sequence based on said determined correspondences, said determining motion using a modified random sample consensus algorithm that selects samples from buckets, iterates a model estimation multiple times including all inliers obtained so far in each iteration, and finds a model that maximizes a data likelihood, wherein a motion hypothesis is derived with a least squares method when the obtained inliers are determined to be less than a number, and the motion hypothesis is derived with a weighted total least squares method otherwise;
  
  generating a forward warping matrix and a background warping matrix for each of said frames based on said determined motion;
  
  generating a forward warping error and a backward warping error for each of said frames based on said forward warping matrix and said background warping matrix;
  
  generating a foreground/background mask for each of said frames based on said forward warping error and said backward warping error; and
  
  generating a background mosaic by mapping said frames to a common coordinate system; and
  
  extracting foreground information from each of said frames based on said background mosaic.
- View Dependent Claims (2, 3, 4, 5, 16, 17, 18)
- - 2. The method of claim 1, wherein said feature extracting step includes extracting a scale invariant feature transform feature.
  - 3. The method of claim 1, wherein said determining correspondences step includes checking for temporal consistency of said extracted feature along said frames.
  - 4. The method of claim 1, wherein said generating a background mosaic step further includes inpainting missing regions of said video sequence from surrounding regions of said video sequence.
  - 5. The method of claim 1, wherein said foreground information extracting step includes extracting said foreground information using a mean shift method.
  - 16. The method of claim 1, wherein the modified random sample consensus algorithm iterates until a constraint parameter is met.
  - 17. The method of claim 16, wherein the constraint parameter is relaxed after each iteration.
  - 18. The method of claim 16, wherein the constraint parameter is a threshold value for determining when a data point fits a model.

6. An apparatus for processing a video sequence comprised of a plurality of frames, said apparatus comprising:
- a processor configured to extract a feature from each of said frames, to determine correspondences between said extracted feature from said frames, to determine motion in said video sequence based on said determined correspondences, said determination of motion being performed using a modified random sample consensus algorithm that selects samples from buckets, iterates a model estimation multiple times including all inliers obtained so far in each iteration, and finds a model that maximizes a data likelihood, wherein a motion hypothesis is derived with a least squares method when the obtained inliers are determined to be less than a number, and the motion hypothesis is derived with a weighted total least squares method otherwise, to generate a forward warping matrix and a background warping matrix for each of said frames based on said determined motion, to generate a forward warping error and a backward warping error for each of said frames based on said forward warping matrix and said background warping matrix, to generate a foreground/background mask for each of said frames based on said forward warping error and said backward warping error, and to generate a background mosaic by mapping said frames to a common coordinate system, and to extract foreground information from each of said frames based on said background mosaic.
- View Dependent Claims (7, 8, 9, 10)
- - 7. The apparatus of claim 6, wherein said processor is further configured to extract a scale invariant feature transform feature.
  - 8. The apparatus of claim 6, wherein said processor is further configured to check for temporal consistency of said extracted feature along said frames.
  - 9. The apparatus of claim 6, wherein said processor is further configured to inpaint missing regions of said video sequence from surrounding regions of said video sequence.
  - 10. The apparatus of claim 6, wherein said processor is further configured to extract said foreground information using a mean shift method.

11. An apparatus for processing a video sequence comprised of a plurality of frames, said apparatus comprising circuitry configured to perform:
- extracting a feature from each of said frames;
  
  determining correspondences between said extracted features from said frames;
  
  determining motion in said video sequence based on said determined correspondences, said determining using a modified random sample consensus algorithm that selects samples from buckets, iterates a model estimation multiple times including all inliers obtained so far in each iteration, and finds a model that maximizes a data likelihood, wherein a motion hypothesis is derived with a least squares method when the obtained inliers are determined to be less than a number, and the motion hypothesis is derived with a weighted total least squares method otherwise;
  
  generating a forward warping matrix and a background warping matrix for each of said frames based on said determined motion;
  
  generating a forward warping error and a backward warping error for each of said frames based on said forward warping matrix and said background warping matrix;
  
  generating a foreground/background mask for each of said frames based on said forward warping error and said backward warping error; and
  
  generating a background mosaic by mapping said frames to a common coordinate system; and
  
  extracting foreground information from each of said frames based on said background mosaic.
- View Dependent Claims (12, 13, 14, 15)
- - 12. The apparatus of claim 11, wherein said feature extracting step includes extracting a scale invariant feature transform feature.
  - 13. The apparatus of claim 11, wherein said determining correspondences step includes checking for temporal consistency of said extracted features along said frames.
  - 14. The apparatus of claim 11, wherein said generating a background mosaic step further includes inpainting missing regions of said video sequence from surrounding regions of said video sequence.
  - 15. The apparatus of claim 11, wherein said foreground information extracting step includes extracting said foreground information using a mean shift method.

Specification

Resources

Litigation Campaign Assessment

Current Assignee
InterDigital Madison Patent Holdings (InterDigital, Inc.)
Original Assignee
Thomson Licensing (Vantiva SA)
Inventors
Sole, Joel, Huang, Yu, Llach, Joan
Primary Examiner(s)
Kelley, Christopher S
Assistant Examiner(s)
Retallick, Kaitlin A

Application Number

US12/451,264
Publication Number

US 20100128789A1
Time in Patent Office

2,790 Days
Field of Search

None
US Class Current

1/1
CPC Class Codes

G06T 2207/10016   Video; Image sequence

G06T 5/77   Retouching; Inpainting; Scr...

G06T 7/194   involving foreground-backgr...

G06T 7/215   Motion-based segmentation

Method and apparatus for processing video sequences

First Claim

4 Assignments

0 Petitions

Accused Products

Abstract

33 Citations

18 Claims

Specification

Solutions

Use Cases

Quick Links

Method and apparatus for processing video sequences

First Claim

4 Assignments

Subscription Required

Subscription Required

0 Petitions

Subscription Required

Accused Products

Subscription Required

Abstract

33 Citations

18 Claims

Specification

Subscription Required

Solutions

Use Cases

Quick Links