For buyers
Training data, by category
Request a sample to see structure, formats and licence terms for a specific collection.
Dataset categories
Text corpora
Written material for pretraining and fine-tuning, from general text to specialist domains.
- modality
- text, code
- format
- JSONL, Parquet, plain text
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
Human-labeled and RLHF data
Annotations, preference rankings and expert-written demonstrations.
- modality
- text, image, audio
- format
- JSONL, CSV
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
Multimodal data
Images, audio and video paired with captions, transcripts or metadata.
- modality
- image, audio, video, text
- format
- Parquet, WebDataset, original media files
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
Synthetic data
Generated datasets that are curated and filtered for training use.
- modality
- text, code, tabular
- format
- JSONL, Parquet, CSV
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
Other training data
Sensor, tabular, geospatial, scientific and other specialist data.
- modality
- tabular, geospatial, time series, scientific
- format
- CSV, Parquet, domain-specific formats
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range