VidChapters-7M: Video Chapters at Scale

The result's identifiers

Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21730%2F23%3A00372038" target="_blank" >RIV/68407700:21730/23:00372038 - isvavai.cz</a>
Result on the web
<a href="https://nips.cc/virtual/2023/poster/73545" target="_blank" >https://nips.cc/virtual/2023/poster/73545</a>
DOI - Digital Object Identifier
—

Alternative languages

Result language
angličtina
Original language name
VidChapters-7M: Video Chapters at Scale
Original language description
Segmenting long videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly released datasets. To address this issue, we present VidChapters-7M, a dataset of 817K user-chaptered videos including 7M chapters in total. VidChapters7M is automatically created from videos online in a scalable manner by scraping user-annotated chapters and hence without any additional manual annotation. We introduce the following three tasks based on this data. First, the video chapter generation task consists of temporally segmenting the video and generating a chapter title for each segment. To further dissect the problem, we also define two variants of this task: video chapter generation given ground-truth boundaries, which requires generating a chapter title given an annotated video segment, and video chapter grounding, which requires temporally localizing a chapter given its annotated title. We benchmark both simple baselines and state-of-the-art video-language models for these three tasks. We also show that pretraining on VidChapters-7M transfers well to dense video captioning tasks in both zero-shot and finetuning settings, largely improving the state of the art on the YouCook2 and ViTT benchmarks. Finally, our experiments reveal that downstream performance scales well with the size of the pretraining dataset. Our dataset, code, and models are publicly available at https://antoyang.github.io/vidchapters.html.
Czech name
—
Czech description
—

Classification

Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

Project
<a href="/en/project/EF15_003%2F0000468" target="_blank" >EF15_003/0000468: Intelligent Machine Perception</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

Publication year
2023
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

Article name in the collection
Advances in Neural Information Processing Systems 36 (NeurIPS 2023)
ISBN
—
ISSN
1049-5258
e-ISSN
—
Number of pages
17
Pages from-to
49428-49444
Publisher name
Neural Information Processing Society
Place of publication
Montreal
Event location
New Orleans
Event date
Dec 10, 2023
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
001230083404009

Similar results(10)

Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning TubeDETR: Spatio-Temporal Video Grounding with Transformers Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains

What are you looking for?

Quick search

Smart search

VidChapters-7M: Video Chapters at Scale

The result's identifiers

Alternative languages

Classification

Result continuities

Others

Data specific for result type

Similar results(10)

What are you looking for?

Quick search

Smart search

Result description

The result's identifiers

The result's identifiers

Alternative languages

Alternative languages

Classification

Classification

Result continuities

Result continuities

Others

Others

Data specific for result type

Data specific for result type

Similar results(10)