Dataset category
Multimodal data
Images, audio and video paired with captions, transcripts or metadata.
Image, audio and video collections with the text that goes with them, such as captions, transcripts and descriptive metadata. Request a sample to see formats and license terms for a specific collection.
- modality
- image, audio, video, text
- format
- Parquet, WebDataset, original media files
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
- example uses
- Vision-language model training
- Speech recognition and synthesis
- Video understanding