Classifier tuning based on data similarities

US 7,089,241 B1
Filed: 12/22/2003
Issued: 08/08/2006
Est. Priority Date: 01/24/2003
Status: Expired due to Fees

First Claim

Patent Images

1. A machine readable medium storing one or more programs that implement an e-mail classifier for determining whether at least one received e-mail should be classified as spam, the one or more programs comprising instructions for causing one or more processing devices to perform the following operations:

obtain feature data for the received e-mail by determining whether the received e-mail has a predefined set of features;

train a scoring classifier using a set of unique training e-mails;

provide a classification output, using the scoring classifier, based on the obtained feature data, wherein the classification output is indicative of whether or not the received e-mail is spam;

compare the provided classification output to a classification threshold, wherein the received e-mail is classified as spam when the comparison of the classification output to the classification threshold indicates the received e-mail is spam;

determine at least one similarity rate for at least one e-mail, wherein the at least one similarity rate is the rate at which e-mails, which are substantially similar to the at least one e-mail, are received by the e-mail classifier;

select and set a value for the classification threshold, wherein selecting and setting the value for the classification threshold includes;

selecting and setting an initial value for the classification threshold that reduces misclassification costs based on a set of unique evaluation e-mails; and

selecting and setting a new value for the classification threshold that reduces the misclassification costs based at least on the determined at least one similarity rate.

View all claims

10 Assignments

Timeline View

Assignment View

0 Petitions

Accused Products

Abstract

A probabilistic classifier is used to classify data items in a data stream. The probabilistic classifier is trained, and an initial classification threshold is set, using unique training and evaluation data sets (i.e., data sets that do not contain duplicate data items). Unique data sets are used for training and in setting the initial classification threshold so as to prevent the classifier from being improperly biased as a result of similarity rates in the training and evaluation data sets that do not reflect similarity rates encountered during operation. During operation, information regarding the actual similarity rates of data items in the data stream is obtained and used to adjust the classification threshold such that misclassification costs are minimized given the actual similarity rates.

Citations

27 Claims

1. A machine readable medium storing one or more programs that implement an e-mail classifier for determining whether at least one received e-mail should be classified as spam, the one or more programs comprising instructions for causing one or more processing devices to perform the following operations:
- obtain feature data for the received e-mail by determining whether the received e-mail has a predefined set of features;
  
  train a scoring classifier using a set of unique training e-mails;
  
  provide a classification output, using the scoring classifier, based on the obtained feature data, wherein the classification output is indicative of whether or not the received e-mail is spam;
  
  compare the provided classification output to a classification threshold, wherein the received e-mail is classified as spam when the comparison of the classification output to the classification threshold indicates the received e-mail is spam;
  
  determine at least one similarity rate for at least one e-mail, wherein the at least one similarity rate is the rate at which e-mails, which are substantially similar to the at least one e-mail, are received by the e-mail classifier;
  
  select and set a value for the classification threshold, wherein selecting and setting the value for the classification threshold includes;
  
  selecting and setting an initial value for the classification threshold that reduces misclassification costs based on a set of unique evaluation e-mails; and
  
  selecting and setting a new value for the classification threshold that reduces the misclassification costs based at least on the determined at least one similarity rate.
- View Dependent Claims (2, 3, 4, 5, 6, 7, 8, 9)
- - 2. The medium of claim 1 wherein the new value for the classification threshold is selected and set also based on the classification output for the at least one e-mail.
  - 3. The medium of claim 1 wherein the new value for the classification threshold is selected and set also based on a classification indication for the at least one e-mail.
  - 4. The medium of claim 1 wherein the initial value of the classification threshold minimizes the misclassification costs.
  - 5. The medium of claim 1 wherein the new value of the classification threshold minimizes the misclassification costs.
  - 6. The medium of claim 1 wherein the misclassification costs depend on varying costs of misclassifying subcategories of non-spam e-mail as spam e-mail.
  - 7. The medium of claim 1 wherein, to select and set the initial value for the classification threshold, the one or more programs comprise instructions for causing the one or more processing devices to perform the following operations:
    - obtain known classes for unique evaluation e-mails in the set of unique evaluation e-mails;
      
      obtain classification outputs indicative of whether or not the unique evaluation emails in the set of unique evaluation e-mails belong to a particular class; and
      
      determine the initial value of the classification threshold based on the classification outputs and the known classes.
  - 8. The medium of claim 7 wherein the initial value of the classification threshold minimizes the misclassification costs.
  - 9. The medium of claim 8 wherein the misclassification costs depend on varying costs of misclassifying subcategories of non-spam e-mail as spam e-mail.

10. A method for determining whether at least one received e-mail should be classified as spam, the method comprising:
- obtaining feature data for the received e-mail by determining whether the received e-mail has a predefined set of features;
  
  training a scoring classifier using a set of unique training e-mails;
  
  providing a classification output, using the scoring classifier, based on the obtained feature data, wherein the classification output is indicative of whether or not the received e-mail is spam;
  
  comparing the provided classification output to a classification threshold, wherein the received e-mail is classified as spam when the comparison of the classification output to the classification threshold indicates the received e-mail is spam;
  
  determining at least one similarity rate for at least one e-mail, wherein the at least one similarity rate is the rate at which e-mails, which are substantially similar to the at least one e-mail, are received by an e-mail classifier;
  
  selecting and setting a value for the classification threshold, wherein selecting and setting the value for the classification threshold includes;
  
  selecting and setting an initial value for the classification threshold that reduces misclassification costs based on a set of unique evaluation e-mails; and
  
  selecting and setting a new value for the classification threshold that reduces the misclassification costs based at least on the determined at least one similarity rate.
- View Dependent Claims (11, 12, 13, 14, 15, 16, 17, 18)
- - 11. The method of claim 10 wherein the new value for the classification threshold is selected and set also based on the classification output for the at least one e-mail.
  - 12. The method of claim 10 wherein the new value for the classification threshold is selected and set also based on a classification indication for the at least one e-mail.
  - 13. The method of claim 10 wherein the initial value of the classification threshold minimizes misclassification costs.
  - 14. The method of claim 10 wherein the new value of the classification threshold minimizes misclassification costs.
  - 15. The method claim 10 wherein the misclassification costs depend on varying costs of misclassifying subcategories of non-spam e-mail as spam e-mail.
  - 16. The method of claim 10 wherein selecting and setting the initial value for the classification threshold comprises:
    - obtaining known classes for unique evaluation e-mails in the set of unique evaluation e-mails;
      
      obtaining classification outputs indicative of whether or not the unique evaluation emails in the set of unique evaluation e-mails belong to a particular class; and
      
      determining the initial value of the classification threshold based on the classification outputs and the known classes.
  - 17. The method of claim 16 wherein the initial value of the classification threshold minimizes the misclassification costs.
  - 18. The method of claim 17 wherein the misclassification costs depend on varying costs of misclassifying subcategories of non-spam e-mail as spam e-mail.

19. An e-mail server that determines whether at least one received e-mail should be classified as spam, the e-mail server comprising:
- one or more processing devices configured to implement the following operations;
  
  obtain feature data for the received e-mail by determining whether the received e-mail has a predefined set of features;
  
  train a scoring classifier using a set of unique training e-mails;
  
  provide a classification output, using the scoring classifier, based on the obtained feature data, wherein the classification output is indicative of whether or not the received e-mail is spam;
  
  compare the provided classification output to a classification threshold, wherein the received e-mail is classified as spam when the comparison of the classification output to the classification threshold indicates the received e-mail is spam;
  
  determine at least one similarity rate for at least one e-mail, wherein the at least one similarity rate is the rate at which e-mails, which are substantially similar to the at least one e-mail, are received by an e-mail classifier;
  
  select and set a value for the classification threshold, wherein selecting and setting the value for the classification threshold includes;
  
  selecting and setting an initial value for the classification threshold that reduces misclassification costs based on a set of unique evaluation e-mails; and
  
  selecting and setting a new value for the classification threshold that reduces the misclassification costs based at least on the determined at least one similarity rate.
- View Dependent Claims (20, 21, 22, 23, 24, 25, 26, 27)
- - 20. The e-mail server of claim 19 wherein the new value for the classification threshold is selected and set also based on the classification output for the at least one e-mail.
  - 21. The e-mail server of claim 19 wherein the new value for the classification threshold is selected and set also based on a classification indication for the at least one e-mail.
  - 22. The e-mail server of claim 19 wherein the initial value of the classification threshold minimizes misclassification costs.
  - 23. The e-mail server of claim 19 wherein the new value of the classification threshold minimizes misclassification costs.
  - 24. The e-mail server of claim 19 wherein the misclassification costs depend on varying costs of misclassifying subcategories of non-spam e-mail as spam e-mail.
  - 25. The e-mail server of claim 19 wherein, to select and set the initial value for the classification threshold, the one or more programs include instructions for causing the one or more processing devices to perform the following operations:
    - obtain known classes for unique evaluation e-mails in the set of unique evaluation e-mails;
      
      obtain classification outputs indicative of whether or not the unique evaluation e-mails in the set of unique evaluation e-mails belong to a particular class; and
      
      determine the initial value of the classification threshold based on the classification outputs and the known class.
  - 26. The e-mail server of claim 25 wherein the initial value of the classification threshold minimizes the misclassification costs.
  - 27. The e-mail server of claim 26 wherein the misclassification costs depend on varying costs of misclassifying subcategories of non-spam e-mail as spam e-mail.

Specification

Resources

Litigation Campaign Assessment

Current Assignee
Google LLC (Alphabet Inc.)
Original Assignee
America Online Inc. (Warner Bros. Discovery, Inc.)
Inventors
Chowdhury, Abdur, Kolcz, Aleksander, Alspector, Joshua
Primary Examiner(s)
Gaffin, Jeffrey A.
Assistant Examiner(s)
Pham, Hung

Application Number

US10/740,821
Time in Patent Office

960 Days
Field of Search

707/7, 707/10, 707/104.1, 709/205, 709/206, 709/207
US Class Current

1/1
CPC Class Codes

G06N 20/00   Machine learning

G06Q 10/107   Computer-aided management o...

H04L 51/212   using filtering or selectiv...

Y10S 707/99937   Sorting

Y10S 707/99945   Object-oriented database st...

Classifier tuning based on data similarities

First Claim

10 Assignments

0 Petitions

Accused Products

Abstract

Citations

27 Claims

Specification

Solutions

Use Cases

Quick Links

Classifier tuning based on data similarities

First Claim

10 Assignments

Subscription Required

Subscription Required

0 Petitions

Subscription Required

Accused Products

Subscription Required

Abstract

Citations

27 Claims

Specification

Subscription Required

Solutions

Use Cases

Quick Links