Text Classification Using ESC-based Stochastic Decision Lists (2002)
Cached
Download Links
- [research.microsoft.com]
- [www.research.microsoft.com]
- DBLP
Other Repositories/Bibliography
| Venue: | In Proceedings of CIKM-99, 8th ACM International Conference on Information and Knowledge Management |
| Citations: | 12 - 3 self |
BibTeX
@INPROCEEDINGS{Li02textclassification,
author = {Hang Li and Kenji Yamanishi},
title = {Text Classification Using ESC-based Stochastic Decision Lists},
booktitle = {In Proceedings of CIKM-99, 8th ACM International Conference on Information and Knowledge Management},
year = {2002},
pages = {122--130}
}
OpenURL
Abstract
We propose a new method of text classification using stochastic decision lists. A stochastic decision list is an ordered sequence of IF-THEN rules, and our method can be viewed as a rule-based method for text classification having advantages of readability and refinability of acquired knowledge. Our method is unique in that decision lists are automatically constructed on the basis of the principle of minimizing Extended Stochastic Complexity (ESC), and with it we are able to construct decision lists that have fewer errors in classification. The accuracy of classification achieved with our method appears better than or comparable to those of existing rule-based methods. We have empirically demonstrated that rule-based methods like ours result in high classification accuracy when the categories to which texts are to be assigned are relatively specific ones and when the texts tend to be short. We have also empirically verified the advantages of rule-based methods over non-rule-based ones.







