Statistical Hypothesis Testing in Positive Unlabelled Data
Abstract
We propose a set of novel methodologies which enable valid statistical hypothesis testing when we have only positive and unlabelled (PU) examples. This type of problem, a special case of semi-supervised data, is common in text mining, bioinformatics, and computer vision. Focusing on a generalised likelihood ratio test, we have 3 key contributions: (1) a proof that assuming all unlabelled examples are negative cases is sufficient for independence testing, but not for power analysis activities; (2) a new methodology that compensates this and enables power analysis, allowing sample size determination for observing an effect with a desired power; and finally, (3) a new capability, supervision determination , which can determine a-priori the number of labelled examples the user must collect before being able to observe a desired statistical effect. Beyond general hypothesis testing, we suggest the tools will additionally be useful for information theoretic feature selection, and Bayesian Network structure learning.
Cite
Text
Sechidis et al. "Statistical Hypothesis Testing in Positive Unlabelled Data." European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, 2014. doi:10.1007/978-3-662-44845-8_5Markdown
[Sechidis et al. "Statistical Hypothesis Testing in Positive Unlabelled Data." European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, 2014.](https://mlanthology.org/ecmlpkdd/2014/sechidis2014ecmlpkdd-statistical/) doi:10.1007/978-3-662-44845-8_5BibTeX
@inproceedings{sechidis2014ecmlpkdd-statistical,
title = {{Statistical Hypothesis Testing in Positive Unlabelled Data}},
author = {Sechidis, Konstantinos and Calvo, Borja and Brown, Gavin},
booktitle = {European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases},
year = {2014},
pages = {66-81},
doi = {10.1007/978-3-662-44845-8_5},
url = {https://mlanthology.org/ecmlpkdd/2014/sechidis2014ecmlpkdd-statistical/}
}