PHRASE-BASED DETECTION OF DUPLICATE DOCUMENTS IN AN INFORMATION RETRIEVAL SYSTEM
First Claim
Patent Images
1. A method of detecting duplicate documents in search results, the method comprising:
- receiving a query comprising at least one phrase;
retrieving a plurality of documents responsive to the query to form a search result;
for each of the retrieved documents, generating a document description comprising selected sentences of the document, wherein the selected sentences are ordered in the document description as a function of a number of phrases in each sentence;
responsive to the document description at least two documents matching, discarding at least one of the two documents from the search result.
2 Assignments
0 Petitions
Accused Products
Abstract
An information retrieval system uses phrases to index, retrieve, organize and describe documents. Phrases are identified that predict the presence of other phrases in documents. Documents are the indexed according to their included phrases. Related phrases and phrase extensions are also identified. Phrases in a query are identified and used to retrieve and rank documents. Phrases are also used to cluster documents in the search results, create document descriptions, and eliminate duplicate documents from the search results, and from the index.
-
Citations
2 Claims
-
1. A method of detecting duplicate documents in search results, the method comprising:
-
receiving a query comprising at least one phrase; retrieving a plurality of documents responsive to the query to form a search result; for each of the retrieved documents, generating a document description comprising selected sentences of the document, wherein the selected sentences are ordered in the document description as a function of a number of phrases in each sentence; responsive to the document description at least two documents matching, discarding at least one of the two documents from the search result.
-
-
2. A method of detecting duplicate documents in search results, the method comprising:
-
receiving a query comprising at least one phrase; retrieving a plurality of documents responsive to the query to form a search result; for each of the retrieved documents, retrieving a stored document description comprising selected sentences of the document, wherein the selected sentences are ordered in the document description as a function of a number of phrases in each sentence; responsive to the document description at least two documents matching, discarding at least one of the two documents from the search result.
-
Specification