×

Intelligent system that dynamically improves its knowledge and code-base for natural language understanding

  • US 9,965,458 B2
  • Filed: 12/09/2015
  • Issued: 05/08/2018
  • Est. Priority Date: 12/09/2014
  • Status: Active Grant
First Claim
Patent Images

1. A method for tokenizing text for natural language processing, the method comprising:

  • generating, by one or more processors in a natural language processing platform, and from a pool of documents, a set of statistical models comprising one or more entries each indicating a likelihood of appearance of a character/letter sequence in the pool of documents;

    receiving, by the one or more processors, a set of rules comprising rules that identify character/letter sequences as valid tokens;

    transforming, by the one or more processors, one or more entries in the statistical models into new rules that are added to the set of rules when the entries indicate a high likelihood;

    receiving, by the one or more processors, a document to be processed;

    dividing, by the one or more processors, the document to be processed into tokens based on the set of statistical models and the set of rules, wherein the statistical models are applied where the rules fail to unambiguously tokenize the document; and

    outputting, by the one or more processors, the divided tokens for natural language processing.

View all claims
  • 13 Assignments
Timeline View
Assignment View
    ×
    ×