LAION has introduced LAION-BVD (LAION — Big Video Dataset), an open-access video dataset for multimodal learning. The dataset comprises 1.3 billion platform-specific video URLs collected from CommonCrawl, from which 80 million videos, amounting to 10 million hours of content, have been downloaded and processed.
LAION-BVD is structured to facilitate multimodal pre-training across video, audio, and image modalities. The dataset uses content-aware scene detection to extract clips, for which video and audio captions are synthetically generated. It also extracts scene-changing frames to serve as an alternative source for image-text data.
Models trained using LAION-BVD have demonstrated competitive performance on standard video-text and audio-text benchmarks. For instance, ViCLIP models trained on LAION-BVD matched or exceeded InternVid-trained models by up to 2.1% on video-text benchmarks. CLAP models achieved competitive results against other large-scale uncurated audio datasets, and frame-based CLIP models showed strong image-text retrieval performance.
The release of LAION-BVD aims to support open and reproducible multimodal research. This initiative addresses the concentration of large-scale video datasets and trained models within proprietary technology companies, providing an open resource for independent scientific investigation and reproducibility in the academic research community.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
LAION has released LAION-BVD, a large-scale open video dataset containing 1.3 billion platform-specific video URLs, from which 80 million videos totaling 10 million hours were downloaded. This dataset is designed to support multimodal pre-training across video, audio, and image modalities, providing an open resource for academic research in AI.