Systems and methods for predictive coding

US 8,554,716 B1
Filed: 03/28/2011
Issued: 10/08/2013
Est. Priority Date: 05/25/2010
Status: Active Grant

First Claim

Patent Images

1. A method for analyzing a plurality of documents, comprising:

receiving the plurality of documents via a computing device;

receiving user input from the computing device, the user input including hard coding of a subset of the plurality of documents, the hard coding based on an identified subject or category; and

executing instructions stored in memory, wherein execution of the instructions by a processor;

generates an initial control set based on the subset of the plurality of documents and the received user input on the subset,analyzes the initial control set to determine at least one seed set parameter associated with the identified subject or category,automatically codes a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category,analyzes the first portion of the plurality of documents by applying an adaptive identification cycle, the adaptive identification cycle being based on the initial control set, user validation of the automated coding of the first portion of the plurality of documents and confidence threshold validation, andretrieves a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle on the first portion of the plurality of documents.

View all claims

10 Assignments

Timeline View

Assignment View

0 Petitions

Accused Products

Abstract

Systems and methods for analyzing documents are provided herein. A plurality of documents and user input are received via a computing device. The user input includes hard coding of a subset of the plurality of documents, based on an identified subject or category. Instructions stored in memory are executed by a processor to generate an initial control set, analyze the initial control set to determine at least one seed set parameter, automatically code a first portion of the plurality of documents based on the initial control set and the seed set parameter associated with the identified subject or category, analyze the first portion of the plurality of documents by applying an adaptive identification cycle, and retrieve a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle test on the first portion of the plurality of documents.

41 Citations

View as Search Results

20 Claims

1. A method for analyzing a plurality of documents, comprising:
- receiving the plurality of documents via a computing device;
  
  receiving user input from the computing device, the user input including hard coding of a subset of the plurality of documents, the hard coding based on an identified subject or category; and
  
  executing instructions stored in memory, wherein execution of the instructions by a processor;
  
  generates an initial control set based on the subset of the plurality of documents and the received user input on the subset,analyzes the initial control set to determine at least one seed set parameter associated with the identified subject or category,automatically codes a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category,analyzes the first portion of the plurality of documents by applying an adaptive identification cycle, the adaptive identification cycle being based on the initial control set, user validation of the automated coding of the first portion of the plurality of documents and confidence threshold validation, andretrieves a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle on the first portion of the plurality of documents.

2. A method for analyzing a plurality of documents, comprising:
- receiving the plurality of documents via a computing device;
  
  filtering the plurality of documents to produce a subset of the plurality of documents;
  
  receiving user input from the computing device, the user input including hard coding of the subset of the plurality of documents, the hard coding based on an identified subject or category;
  
  executing instructions stored in memory, wherein execution of the instructions by a processor;
  
  generates an initial control set based on random sampling of the subset of the plurality of documents on a static basis and a rolling basis,reviews the initial control set to determine at least one seed set parameter associated with the identified subject or category, andautomatically codes a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category;
  
  receiving user input via the computing device, the user input comprising inspection, analysis and coding of initial control set documents; and
  
  executing instructions stored in memory, wherein execution of the instructions by the processor automatically codes documents based on the received user input regarding the initial control set documents.
- View Dependent Claims (3, 4, 6, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18)
- - 3. The method of claim 2, wherein the filtering further comprises culling the plurality of documents based on metadata.
  - 4. The method of claim 2, further comprising receiving further user input from the computing device, the further user input comprising a designation corresponding to key documents of the initial control set, the key documents comprising documents likely to be responsive to a coding category.
  - 6. The method of claim 2, further comprising executing further instructions stored in memory, wherein execution of the further instructions by the processor adds further documents to the plurality of documents on a rolling load basis.
  - 8. The method of claim 2, further comprising transmitting to a display the first portion of the plurality of documents, the display coupled to the computing device.
  - 9. The method of claim 2, further comprising:
    - receiving user input via the computing device, the user input corresponding to a confidence level; and
      
      executing further instructions stored in memory, wherein execution of the further instructions by the processor;
      
      calculates a statistic regarding machine-only coding accuracy rate, andcompares the statistic regarding machine-only coding accuracy rate against user input based on a defined confidence interval.
  - 10. The method of claim 2, wherein the step of automatically coding the first portion of the plurality of documents further comprises executing instructions stored in memory, wherein execution of the instructions by the processor automatically codes based on probabilistic latent semantic analysis and support vector machine analysis of the first portion of the plurality of documents.
  - 11. The method of claim 2, further comprising applying targeted document identification to the plurality of documents.
  - 12. The method of claim 11, wherein applying further comprises identifying a candidate key document based on one or more search terms.
  - 13. The method of claim 11, wherein applying further comprises identifying a candidate key document based on sampling based on one or more phrases.
  - 14. The method of claim 11, wherein applying further comprises identifying a candidate key document using a location filter.
  - 15. The method of claim 11, wherein applying further comprises identifying a candidate key document based on sampling and extrapolation.
  - 16. The method of claim 2, further comprising applying confidence threshold validation testing on one or more of the plurality of documents.
  - 17. The method of claim 16, wherein applying confidence threshold validation testing comprises:
    - setting a size of a quality control (QC) sample set at the size of the initial control set;
      
      creating the QC sample set by random sampling from unreviewed documents; and
      
      reviewing the QC sample set.
  - 18. The method of claim 2, further comprising adding documents on a rolling basis for automatic coding based on received user input regarding the initial control set documents.

5. A method for analyzing a plurality of documents, comprising:
- receiving the plurality of documents via a computing device;
  
  filtering the plurality of documents to produce a subset of the plurality of documents;
  
  receiving user input from the computing device, the user input including hard coding of the subset of the plurality of documents, the hard coding based on an identified subject or category;
  
  executing instructions stored in memory, wherein execution of the instructions by a processor;
  
  generates an initial control set based on random sampling of the subset of the plurality of documents on a static basis and a rolling basis,reviews the initial control set to determine at least one seed set parameter associated with the identified subject or category, andautomatically codes a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category;
  
  receiving user input via the computing device, the user input comprising inspection, analysis and coding of initial control set documents; and
  
  executing instructions stored in memory, wherein execution of the instructions by the processor;
  
  automatically codes documents based on the received user input regarding the initial control set documents;
  
  analyzes the first portion of the plurality of documents by applying an adaptive identification cycle, the adaptive identification cycle being based on the initial control set, user validation of the automated coding of the first portion of the plurality of documents and confidence threshold validation, andretrieves a second portion of the plurality of documents based on a result of the application of the adaptive identification cycle on the first portion of the plurality of documents.

7. A method for analyzing a plurality of documents, comprising:
- receiving the plurality of documents via a computing device;
  
  filtering the plurality of documents to produce a subset of the plurality of documents;
  
  receiving user input from the computing device, the user input including hard coding of the subset of the plurality of documents, the hard coding based on an identified subject or category;
  
  executing instructions stored in memory, wherein execution of the instructions by a processor;
  
  generates an initial control set based on random sampling of the subset of the plurality of documents on a static basis and a rolling basis,reviews the initial control set to determine at least one seed set parameter associated with the identified subject or category, andautomatically codes a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category;
  
  receiving user input via the computing device, the user input comprising inspection, analysis and coding of initial control set documents; and
  
  executing instructions stored in memory, wherein execution of the instructions by the processor;
  
  automatically codes documents based on the received user input regarding the initial control set documents;
  
  automatically codes a second portion of the plurality of documents resulting from an application of user analysis and an adaptive identification cycle; and
  
  adds the coded second portion of the plurality of documents to the initial control set.

19. A method for analyzing a plurality of documents, comprising:
- receiving the plurality of documents via a computing device;
  
  receiving user input from the computing device, the user input including hard coding of a subset of the plurality of documents, the hard coding based on an identified subject or category;
  
  executing instructions stored in memory, wherein execution of the instructions by a processor;
  
  generates an initial control set based on random sampling of the subset of the plurality of documents on a static basis and a rolling basis,determines at least one seed set parameter associated with the identified subject or category in the initial control set, andautomatically codes a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category;
  
  receiving user input via the computing device, the user input comprising coding of initial control set documents; and
  
  executing instructions stored in memory, wherein execution of the instructions by the processor automatically codes a second portion of the plurality of documents based the initial control set documents.

20. A method for analyzing a plurality of documents, comprising:
- receiving the plurality of documents via a computing device;
  
  receiving user input from the computing device, the user input including hard coding of a subset of the plurality of documents, the hard coding based on an identified subject or category;
  
  executing instructions stored in memory, wherein execution of the instructions by a processor;
  
  generates an initial control set based on random sampling of the subset of the plurality of documents on one of a static basis and a rolling basis,identifies at least one seed set parameter associated with the identified subject or category in the initial control set, andautomatically codes a first portion of the plurality of documents, based on the initial control set and the at least one seed set parameter associated with the identified subject or category;
  
  receiving user input via the computing device, the user input comprising manual verification of initial control set documents; and
  
  executing instructions stored in memory, wherein execution of the instructions by the processor automatically codes a second portion of the plurality of documents based on the initial control set documents.

Specification

Resources

Litigation Campaign Assessment

Current Assignee
Open Text Holdings, Inc. (Open Text Corporation)
Original Assignee
Recommind Incorporated (Open Text Corporation)
Inventors
Puzicha, Jan, Vranas, Steve
Primary Examiner(s)
Chaki, Kakali
Assistant Examiner(s)
SITIRICHE, LUIS A

Application Number

US13/074,005
Time in Patent Office

925 Days
Field of Search

None
US Class Current

706/52
CPC Class Codes

G06F 16/93   Document management systems

G06N 20/00   Machine learning

G06N 20/10   using kernel methods, e.g. ...

G06N 5/04   Inference or reasoning models

G06N 5/048   Fuzzy inferencing

G06N 7/01   Probabilistic graphical mod...

Systems and methods for predictive coding

First Claim

10 Assignments

0 Petitions

Accused Products

Abstract

41 Citations

20 Claims

Specification

Solutions

Use Cases

Quick Links

Systems and methods for predictive coding

First Claim

10 Assignments

Subscription Required

Subscription Required

0 Petitions

Subscription Required

Accused Products

Subscription Required

Abstract

41 Citations

20 Claims

Specification

Subscription Required

Solutions

Use Cases

Quick Links