SpamRank - Fully Automatic Link Spam Detection (2005)
| Venue: | In Proceedings of the First International Workshop on Adversarial Information Retrieval on the Web (AIRWeb |
| Citations: | 57 - 4 self |
BibTeX
@INPROCEEDINGS{Benczur05spamrank-,
author = {Andras A. Benczur and Karoly Csalogany and Tamas Sarlos and Mate Uher and Máté Uher},
title = {SpamRank - Fully Automatic Link Spam Detection},
booktitle = {In Proceedings of the First International Workshop on Adversarial Information Retrieval on the Web (AIRWeb},
year = {2005}
}
Years of Citing Articles
OpenURL
Abstract
Spammers intend to increase the PageRank of certain spam pages by creating a large number of links pointing to them. We propose a novel method based on the concept of personalized PageRank that detects pages with an undeserved high PageRank value without the need of any kind of white or blacklists or other means of human intervention. We assume that spammed pages have a biased distribution of pages that contribute to the undeserved high PageRank value. We define SpamRank by penalizing pages that originate a suspicious PageRank share and personalizing PageRank on the penalties. Our method is tested on a 31 M page crawl of the .de domain with a manually classified 1000-page stratified random sample with bias towards large PageRank values.







