Natural language processing system for semantic vector representation which accounts for lexical ambiguity
First Claim
1. A method of generating a subject field code vector representation of a document which comprises the steps of assigning subject codes to each of the words of the document which codes express the semantic content of the document, said codes corresponding to the meanings of each of said words in accordance with the various senses thereof;
- disambiguating said document to select a specific subject code for each of said words heuristically in order first from the occurrence of like codes within each sentence of said documents which occur uniquely and at or with greater than a certain frequency within each sentence, then second correlating the codes for each word with the codes occurring uniquely (unique code) and with greater than or equal to the given frequency in the sentence to select for each word the code having the highest correlation, and then third in accordance with the frequency of usage of the meaning of the word represented by the code; and
arranging said codes into a weighted vector representing the content of said document.
1 Assignment
0 Petitions
Accused Products
Abstract
A natural language processing system uses unformatted naturally occurring text and generates a subject vector representation of the text, which may be an entire document or a part thereof such as its title, a paragraph, clause, or a sentence therein. The subject codes which are used are obtained from a lexical database and the subject code(s) for each word in the text is looked up and assigned from the database. The database may be a dictionary or other word resource which has a semantic classification scheme as designators of subject domains. Various meanings or senses of a word may have assigned thereto multiple, different subject codes and psycholinguistically justified sense meaning disambiguation is used to select the most appropriate subject field code. Preferably, an ordered set of sentence level heuristics is used which is based on the statistical probability or likelihood of one of the plurality of codes being the most appropriate one of the plurality. The subject codes produce a weighted, fixed-length vector (regardless of the length of the document) which represents the semantic content thereof and may be used for various purposes such as information retrieval, categorization of texts, machine translation, document detection, question answering, and generally for extracting knowledge from the document. The system has particular utility in classifying documents by their general subject matter and retrieving documents relevant to a query.
770 Citations
46 Claims
-
1. A method of generating a subject field code vector representation of a document which comprises the steps of assigning subject codes to each of the words of the document which codes express the semantic content of the document, said codes corresponding to the meanings of each of said words in accordance with the various senses thereof;
- disambiguating said document to select a specific subject code for each of said words heuristically in order first from the occurrence of like codes within each sentence of said documents which occur uniquely and at or with greater than a certain frequency within each sentence, then second correlating the codes for each word with the codes occurring uniquely (unique code) and with greater than or equal to the given frequency in the sentence to select for each word the code having the highest correlation, and then third in accordance with the frequency of usage of the meaning of the word represented by the code; and
arranging said codes into a weighted vector representing the content of said document. - View Dependent Claims (2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 41, 43, 44)
- disambiguating said document to select a specific subject code for each of said words heuristically in order first from the occurrence of like codes within each sentence of said documents which occur uniquely and at or with greater than a certain frequency within each sentence, then second correlating the codes for each word with the codes occurring uniquely (unique code) and with greater than or equal to the given frequency in the sentence to select for each word the code having the highest correlation, and then third in accordance with the frequency of usage of the meaning of the word represented by the code; and
-
21. Apparatus for generating a subject field code vector representation of a document which comprises means for assigning subject codes to each of the words of the document which codes express the semantic content of the document, said codes corresponding to the meanings of each of said words in accordance with the various senses thereof;
- means for disambiguating said document to select a specific subject code for each of said words heuristically in order first from the occurrence of like codes within each sentence of said documents which occur uniquely and at or with greater than a certain frequency within each sentence, then second correlating the codes for each word with the codes occurring uniquely (unique code) and with greater than or equal to the given frequency in the sentence to select for each word the code having the highest correlation, and then third in accordance with the frequency of usage of the meaning of the word represented by the code; and
means for arranging said codes into a weighted vector representing the content of said document. - View Dependent Claims (22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 42)
- means for disambiguating said document to select a specific subject code for each of said words heuristically in order first from the occurrence of like codes within each sentence of said documents which occur uniquely and at or with greater than a certain frequency within each sentence, then second correlating the codes for each word with the codes occurring uniquely (unique code) and with greater than or equal to the given frequency in the sentence to select for each word the code having the highest correlation, and then third in accordance with the frequency of usage of the meaning of the word represented by the code; and
-
45. A natural language processing system comprising:
-
means for disambiguating words of a document to provide codes for said words responsive to the meaning and frequency of said words; and means for generating a representation of the document responsive to said codes.
-
-
46. An apparatus for generating a vector representation of a document having on or more sentences comprising:
-
means for selecting subject codes for the words of the document responsive to the meaning and frequency of the words in each said sentence; and means for arranging said codes into the vector representation of the document.
-
Specification