Skip to content
Corpus PrimeRequest a sample

Dataset category

Text corpora

Written material for pretraining and fine-tuning, from general text to specialist domains.

Collections of books, articles, documentation, code, forum threads and domain text in areas such as law, medicine and finance, in many languages. Request a sample to see structure, formats and license terms for a specific collection.

modality
text, code
format
JSONL, Parquet, plain text
license
PLACEHOLDER: license terms
size
PLACEHOLDER: size / volume range
example uses
  • Pretraining and continued pretraining
  • Domain adaptation
  • Retrieval and evaluation sets