ENTITY ANALYSIS SYSTEM
5 Assignments
0 Petitions
Accused Products
Abstract
A method for building a factual database of concepts and entities that are related to the concepts through a learning process. Training content (e.g., news articles, books) and a set of entities (e.g., Bill Clinton and Barack Obama) that are related to a concept (e.g., Presidents) is received. Groups of words that co-occur frequently in the textual content in conjunction with the entities are identified as templates. Templates may also be identified by analyzing parts-of-speech patterns of the templates. Entities that co-occur frequently in the textual content in conjunction with the templates are identified as additional related entities (e.g., Ronald Reagan and Richard Nixon). To eliminate erroneous results, the identified entities may be presented to a user who removes any false positives. The entities are then stored in association with the concept.
-
Citations
41 Claims
-
1-21. -21. (canceled)
-
22. A computer-implemented method of learning related entities, the method comprising:
-
receiving a set of entities, the set of entities including a plurality of entities and each entity in the set of entities relating to a first concept; receiving training content that includes textual content that is organized and that includes the plurality of entities of the set of entities; and learning additional entities that are related to the first concept by iteratively performing the following steps; identifying one or more potential word templates from the training content based on occurrences of one or more words in the training content with an entity of the set of entities, wherein each potential word template is one or more words, and wherein each potential word template is tagged with a part-of-speech tag based on grammatical use of the one or more words ire the training content; identifying one or more word templates from the one or more potential word templates based on a frequency of occurrence of the one or more potential word templates and based on the part-of-speech tag of the one or more potential word templates compared to part-of-speech tags of word templates of a set of word templates, wherein the one or more identified word templates are added to the set of word templates; identifying, for each identified word template, one or more part-of-speech tags of the identified word templates; adjusting, for each identified word template, a confidence score of the identified word template when the one or more part of speech tags of the identified word template is similar to the part-of-speech tags of word templates of a set of word templates; comparing, for each identified word template, the confidence score of the identified word template to a threshold value; and removing the identified word template from the set of word templates when the confidence score of the identified word template is outside the threshold value. - View Dependent Claims (23, 24)
-
-
25. A computer-implemented method of learning related entities, the method comprising:
-
receiving a set of entities, the set of entities including a plurality of entities and each entity in the set of entities relating to a first concept; receiving training content that includes textual content that organized and that includes the plurality of entities of the set of entities; and learning additional entities that are related to the first concept by iteratively performing the following steps; identifying one or more word templates from the training content based on occurrences of one or more words in the training content with an entity of the set of entities, wherein each word template is one or more words, and wherein the one or more identified word templates are added to a set of word templates; identifying, for each identified word template, one or more part-of-speech tags of the identified word templates; adjusting, for each identified word template, a confidence score of the identified word template when the one or more part of speech tags of the identified word template is similar to the part-of-speech tags of word templates of a set of word templates; comparing, for each identified word template, the confidence score of the identified word template to a threshold value; and removing the identified word template from the set of word templates when the confidence score of the identified word template is outside the threshold value. - View Dependent Claims (26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39)
-
-
40. A computer product for learning related entities, the computer product comprising a non-transitory computer-readable medium containing computer program code for performing the method comprising:
-
receiving a set of entities, the set of entities including a plurality of entities and each entity in the set of entities relating to a first concept; receiving training content that includes textual content that is organized and that includes the plurality of entities of the set of entities; and learning additional entities that related to the first concept by iteratively performing the following steps; identifying one or more word templates from the training content based on occurrences of one or more words in the training content with an entity of the set of entities, wherein each word template is one or more words, and wherein the one or more identified word templates are added to a set of word templates; identifying, for each identified word template, one or more part-of-speech tags of the identified word templates; adjusting, for each identified word template, a confidence score of the identified word template when the one or more part of speech tags of the identified word template is similar to the part f-speech tags of word templates of a set of word templates; comparing, for each identified word template, the confidence score of the identified word template to a threshold value; and removing the identified word template from the set of word templates when the confidence score of the identified word template is outside the threshold value. - View Dependent Claims (41)
-
Specification