Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Zhu, Yukun; Kiros, Ryan; Zemel, Rich; Salakhutdinov, Ruslan; Urtasun, Raquel; Torralba, Antonio; Fidler, Sanja

doi:10.1109/ICCV.2015.11

Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, Sanja Fidler

ICCV 2015

doi:10.1109/ICCV.2015.11 /iccv/2015/zhu2015iccv-aligning/

Abstract

Books are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This paper aims to align books to their movie releases in order to provide rich descriptive explanations for visual content that go semantically far beyond the captions available in the current datasets. To align movies and books we propose a neural sentence embedding that is trained in an unsupervised way from a large corpus of books, as well as a video-text neural embedding for computing similarities between movie clips and sentences in the book. We propose a context-aware CNN to combine information from multiple sources. We demonstrate good quantitative performance for movie/book alignment and show several qualitative examples that showcase the diversity of tasks our model can be used for.

PDF ICCV Semantic Scholar

Cite

Text

Zhu et al. "Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books." International Conference on Computer Vision, 2015. doi:10.1109/ICCV.2015.11

Markdown

[Zhu et al. "Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books." International Conference on Computer Vision, 2015.](https://mlanthology.org/iccv/2015/zhu2015iccv-aligning/) doi:10.1109/ICCV.2015.11

BibTeX

@inproceedings{zhu2015iccv-aligning,
  title     = {{Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books}},
  author    = {Zhu, Yukun and Kiros, Ryan and Zemel, Rich and Salakhutdinov, Ruslan and Urtasun, Raquel and Torralba, Antonio and Fidler, Sanja},
  booktitle = {International Conference on Computer Vision},
  year      = {2015},
  doi       = {10.1109/ICCV.2015.11},
  url       = {https://mlanthology.org/iccv/2015/zhu2015iccv-aligning/}
}