A Memory-Based Approach to Anti-Spam Filtering for Mailing Lists (2003)
Cached
Download Links
- [www.aueb.gr]
- [cgi.di.uoa.gr]
- [www.aueb.gr]
- [www.di.uoa.gr]
- [www.aueb.gr]
- [pages.cs.aueb.gr]
- [www.iit.demokritos.gr]
- DBLP
Other Repositories/Bibliography
| Venue: | Information Retrieval |
| Citations: | 34 - 2 self |
BibTeX
@ARTICLE{Sakkis03amemory-based,
author = {Georgios Sakkis and Ion Androutsopoulos and Constantine D. Spyropoulos},
title = {A Memory-Based Approach to Anti-Spam Filtering for Mailing Lists},
journal = {Information Retrieval},
year = {2003},
volume = {6},
pages = {49--73}
}
Years of Citing Articles
OpenURL
Abstract
This paper presents an extensive empirical evaluation of memory-based learning in the context of anti-spam filtering, a novel cost-sensitive application of text categorization that attempts to identify automatically unsolicited commercial messages that flood mailboxes. Focusing on anti-spam filtering for mailing lists, a thorough investigation of the effectiveness of a memory-based anti-spam filter is performed using a publicly available corpus. The investigation includes different attribute and distance-weighting schemes, and studies on the effect of the neighborhood size, the size of the attribute set, and the size of the training corpus. Three different cost scenarios are identified, and suitable cost-sensitive evaluation functions are employed. We conclude that memorybased anti-spam filtering for mailing lists is practically feasible, especially when combined with additional safety nets. Compared to a previously tested Naive Bayes filter, the memory-based filter performs on average better, particularly when the misclassification cost for non-spam messages is high.







