Corpus Prime
Training data with the paperwork to prove it.
We supply training data to AI companies and broker deals between AI companies and the people who own the data.
- record_id
- 000481
- modality
- text
- language
- en
- source
- licensed
- license
- per-agreement
- pii_scrubbed
- true
Every type of training data
Text, human-labeled and RLHF data, multimodal data and synthetic data, plus specialist sources.
Provenance first
Know where data came from and what you may do with it before you train on it.
Brokered deals for data owners
Hold data an AI company would pay for? We find the buyer and do the legwork, and you get paid.
What we supply
Text corpora
Written material for pretraining and fine-tuning, from general text to specialist domains.
- modality
- text, code
- format
- JSONL, Parquet, plain text
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
Human-labeled and RLHF data
Annotations, preference rankings and expert-written demonstrations.
- modality
- text, image, audio
- format
- JSONL, CSV
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
Multimodal data
Images, audio and video paired with captions, transcripts or metadata.
- modality
- image, audio, video, text
- format
- Parquet, WebDataset, original media files
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
Synthetic data
Generated datasets that are curated and filtered for training use.
- modality
- text, code, tabular
- format
- JSONL, Parquet, CSV
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
Other training data
Sensor, tabular, geospatial, scientific and other specialist data.
- modality
- tabular, geospatial, time series, scientific
- format
- CSV, Parquet, domain-specific formats
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range