Dataset-specific open licenses
Dataset collection
Post-Training Dataset Collection
An extensive collection for language-model training with objectives spanning English and Danish instruction and knowledge, mathematics, and agentic-style post-training. The collection contains 86 datasets, with individual datasets reaching up to 2.5 GB of compressed data.
View full collection