Technology
LAION releases 10-million-hour open video dataset for AI research
The LAION-BVD dataset includes 80 million videos sourced from CommonCrawl to facilitate academic AI research across video, audio, and image modalities.
The short version
- LAION released LAION-BVD, an open video dataset containing 80 million downloaded videos totaling 10 million hours, alongside 1.3 billion URLs.
- The dataset is restricted exclusively to non-commercial research to broaden access to large-scale multimodal pre-training data.
- Initial benchmark tests indicate models trained on the dataset perform competitively against existing baselines in video-text, audio-text, and image-text tasks.
Key facts
- LAION-BVD contains 1.3 billion platform-specific video URLs gathered from CommonCrawl, from which 80 million videos totaling 10 million hours were downloaded.[Hacker News]
- The dataset includes synthetic captions for generated video clips as well as 300 million extracted video frames for image-text research.[Hacker News]
- ViCLIP models trained on the dataset outperformed InternVid-trained baselines by up to 2.1 percent on standard video-text benchmarks.[Hacker News]
- LAION released the dataset strictly for scientific and research purposes, prohibiting commercial use.[Hacker News]
What remains uncertain
- LAION acknowledged that the dataset may harbor inherent web biases, stereotypes, uneven demographic or regional representation, and potential copyright issues.[Hacker News]
Sources
- Laion Big Video DatasetHacker News