Dataset category
Text corpora
Written material for pretraining and fine-tuning, from general text to specialist domains.
Collections of books, articles, documentation, code, forum threads and domain text in areas such as law, medicine and finance, in many languages. Request a sample to see structure, formats and license terms for a specific collection.
- modality
- text, code
- format
- JSONL, Parquet, plain text
- license
- PLACEHOLDER: license terms
- size
- PLACEHOLDER: size / volume range
- example uses
- Pretraining and continued pretraining
- Domain adaptation
- Retrieval and evaluation sets