×

Ingestion plan based on table uniqueness

  • US 10,210,193 B2
  • Filed: 06/15/2017
  • Issued: 02/19/2019
  • Est. Priority Date: 11/05/2015
  • Status: Active Grant
First Claim
Patent Images

1. A computer implemented computer program for document processing, the computer program product comprising:

  • one or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media, the program instructions comprising;

    program instructions to receive a plurality of electronic documents on a computer through a network, the plurality of electronic documents being stored on a remote server, the network being an internet connection;

    program instructions to receive a plurality of metadata from the remote server through the network, the metadata being a plurality of identifying information associated with the received plurality of electronic documents;

    program instructions to index the received plurality of electronic documents and the received plurality of metadata in a data store;

    program instructions to identify a plurality of tabular data markers, in response to analyzing the received electronic document and associated metadata;

    program instructions to identify references for association with the identified plurality of tabular data markers by natural language analysis;

    program instructions to generate a graphical representation of the relationship between the identified tabular data markers and identified references, the graphical representation comprising a plurality of inbound directional edges and a plurality of vertices, wherein the directional edges are based on the identified references having an amplitude based on a count of identified references, and the vertices of the plurality of vertices are tabular data of the identified references;

    program instructions to calculate a uniqueness score value based on the generated graphical representation, the uniqueness score comprising a first value based on the plurality of inbound directional edges and a second value based on the plurality of vertices;

    program instructions to modify the calculated uniqueness score based on one or more of;

    a first count based on the directional edges to a vertex in the graphical representation;

    program instructions to multiply the uniqueness score by zero, in response to the count of direction edges not exceeding a threshold; and

    in response to input by a user, a second count based on the of vertexes in the graphical representation exceeding a threshold; and

    program instructions to generate an ingestion plan for the received electronic documents for display based on the calculated uniqueness score value, the ingestion plan comprising an ordered list of the received plurality of electronic documents.

View all claims
  • 1 Assignment
Timeline View
Assignment View
    ×
    ×