Ask the Image: Supervised Pooling to Preserve Feature Locality
Abstract
In this paper we propose a weighted supervised pooling method for visual recognition systems. We combine a standard Spatial Pyramid Representation which is commonly adopted to encode spatial information, with an appropriate Feature Space Representation favoring semantic information in an appropriate feature space. For the latter, we propose a weighted pooling strategy exploiting data supervision to weigh each local descriptor coherently with its likelihood to belong to a given object class. The two representations are then combined adaptively with Multiple Kernel Learning. Experiments on common benchmarks (Caltech-256 and PASCAL VOC-2007) show that our image representation improves the current visual recognition pipeline and it is competitive with similar state-of-art pooling methods. We also evaluate our method on a real Human-Robot Interaction setting, where the pure Spatial Pyramid Representation does not provide sufficient discriminative power, obtaining a remarkable improvement.
Cite
Text
Fanello et al. "Ask the Image: Supervised Pooling to Preserve Feature Locality." Conference on Computer Vision and Pattern Recognition, 2014. doi:10.1109/CVPR.2014.114Markdown
[Fanello et al. "Ask the Image: Supervised Pooling to Preserve Feature Locality." Conference on Computer Vision and Pattern Recognition, 2014.](https://mlanthology.org/cvpr/2014/fanello2014cvpr-ask/) doi:10.1109/CVPR.2014.114BibTeX
@inproceedings{fanello2014cvpr-ask,
title = {{Ask the Image: Supervised Pooling to Preserve Feature Locality}},
author = {Fanello, Sean Ryan and Noceti, Nicoletta and Ciliberto, Carlo and Metta, Giorgio and Odone, Francesca},
booktitle = {Conference on Computer Vision and Pattern Recognition},
year = {2014},
doi = {10.1109/CVPR.2014.114},
url = {https://mlanthology.org/cvpr/2014/fanello2014cvpr-ask/}
}