TV_VTT (TrecVid Video-To-Text) Dataset

Description

This dataset contains short videos (ranging from 3 seconds to 10 seconds) from TRECVID VTT task from 2016 to 2024. There are 73,893 videos with captions. Each video has between 2 and 5 captions, which have been written by dedicated annotators hired by NIST.

Resources

Name Format Description Link
0 https://doi.org/10.18434/mds2-2545
47 A high-level readme file explaining how the dataset (videos and captions) are organized. https://ir.nist.gov/tv_vtt_data/Readme.txt
0 This dataset contains short videos (ranging from 3 seconds to 10 seconds) from TRECVID VTT task from 2016 to 2021. There are 10,862 videos with captions. Each video has between 2 and 5 captions, which have been written by dedicated annotators hired by NIST. https://ir.nist.gov/tv_vtt_data/
47 Please submit the following data agreement form to access the TV_VTT (video to text development dataset) https://data.nist.gov/od/ds/mds2-2545/V3C_VTT_Org.Form.txt

Tags

  • video-retrieval
  • video-to-text
  • image-captioning
  • video-captioning

Topics

Categories