259

Vandalism Detection in Wikipedia: a Bag-of-Words Classifier Approach

Abstract

A bag-of-words based probabilistic classifier is trained using regularized logistic regression to detect vandalism in the English Wikipedia. Isotonic regression is used to calibrate the class membership probabilities. Learning curve, reliability, ROC, and cost analysis are performed.

View on arXiv
Comments on this paper