← Latest briefing

Technology

LAION releases 10-million-hour open video dataset for AI research

The LAION-BVD dataset includes 80 million videos sourced from CommonCrawl to facilitate academic AI research across video, audio, and image modalities.

The short version

  • LAION released LAION-BVD, an open video dataset containing 80 million downloaded videos totaling 10 million hours, alongside 1.3 billion URLs.
  • The dataset is restricted exclusively to non-commercial research to broaden access to large-scale multimodal pre-training data.
  • Initial benchmark tests indicate models trained on the dataset perform competitively against existing baselines in video-text, audio-text, and image-text tasks.

Key facts

  • LAION-BVD contains 1.3 billion platform-specific video URLs gathered from CommonCrawl, from which 80 million videos totaling 10 million hours were downloaded.[Hacker News]
  • The dataset includes synthetic captions for generated video clips as well as 300 million extracted video frames for image-text research.[Hacker News]
  • ViCLIP models trained on the dataset outperformed InternVid-trained baselines by up to 2.1 percent on standard video-text benchmarks.[Hacker News]
  • LAION released the dataset strictly for scientific and research purposes, prohibiting commercial use.[Hacker News]

What remains uncertain

  • LAION acknowledged that the dataset may harbor inherent web biases, stereotypes, uneven demographic or regional representation, and potential copyright issues.[Hacker News]

Sources