Subgroup Discovery with CN2-SD
Abstract
This paper investigates how to adapt standard classification rule learning approaches to subgroup discovery. The goal of subgroup discovery is to find rules describing subsets of the population that are sufficiently large and statistically unusual. The paper presents a subgroup discovery algorithm, CN2-SD, developed by modifying parts of the CN2 classification rule learner: its covering algorithm, search heuristic, probabilistic classification of instances, and evaluation measures. Experimental evaluation of CN2-SD on 23 UCI data sets shows substantial reduction of the number of induced rules, increased rule coverage and rule significance, as well as slight improvements in terms of the area under ROC curve, when compared with the CN2 algorithm. Application of CN2-SD to a large traffic accident data set confirms these findings.
Cite
Text
Lavrač et al. "Subgroup Discovery with CN2-SD." Journal of Machine Learning Research, 2004.Markdown
[Lavrač et al. "Subgroup Discovery with CN2-SD." Journal of Machine Learning Research, 2004.](https://mlanthology.org/jmlr/2004/lavrac2004jmlr-subgroup/)BibTeX
@article{lavrac2004jmlr-subgroup,
title = {{Subgroup Discovery with CN2-SD}},
author = {Lavrač, Nada and Kavšek, Branko and Flach, Peter and Todorovski, Ljupčo},
journal = {Journal of Machine Learning Research},
year = {2004},
pages = {153-188},
volume = {5},
url = {https://mlanthology.org/jmlr/2004/lavrac2004jmlr-subgroup/}
}