Fast and Interpretable Face Identification for Out-of-Distribution Data Using Vision Transformers

Abstract

Most face identification approaches employ a Siamese neural network to compare two images at the image embedding level. Yet, this technique can be subject to occlusion (e.g., faces with masks or sunglasses) and out-of-distribution data. DeepFace-EMD (Phan et al. 2022) reaches state-of-the-art accuracy on out-of-distribution data by first comparing two images at the image level, and then at the patch level. Yet, its later patch-wise re-ranking stage admits a large O(n^3 log n) time complexity (for n patches in an image) due to the optimal transport optimization. In this paper, we propose a novel, 2-image Vision Transformers (ViTs) that compares two images at the patch level using cross-attention. After training on 2M pairs of images on CASIA Webface (Yi et al. 2014), our model performs at a comparable accuracy as DeepFace-EMD on out-of-distribution data, yet at an inference speed more than twice as fast as DeepFace-EMD (Phan et al. 2022). In addition, via a human study, our model shows promising explainability through the visualization of cross-attention. We believe our work can inspire more explorations in using ViTs for face identification.

Cite

Text

Phan et al. "Fast and Interpretable Face Identification for Out-of-Distribution Data Using Vision Transformers." Winter Conference on Applications of Computer Vision, 2024.

Markdown

[Phan et al. "Fast and Interpretable Face Identification for Out-of-Distribution Data Using Vision Transformers." Winter Conference on Applications of Computer Vision, 2024.](https://mlanthology.org/wacv/2024/phan2024wacv-fast/)

BibTeX

@inproceedings{phan2024wacv-fast,
  title     = {{Fast and Interpretable Face Identification for Out-of-Distribution Data Using Vision Transformers}},
  author    = {Phan, Hai and Le, Cindy X. and Le, Vu and He, Yihui and Nguyen, Anh “Totti”},
  booktitle = {Winter Conference on Applications of Computer Vision},
  year      = {2024},
  pages     = {6301-6311},
  url       = {https://mlanthology.org/wacv/2024/phan2024wacv-fast/}
}