News
Extensive open training-data collection on Hugging Face
We published an extensive mix of data with different objectives, ranging from English and Danish instruction and knowledge to mathematics and agentic-style post-training data.
The collection contains 86 datasets, with individual datasets reaching up to 2.5 GB of compressed data. Almost all source data is freely available on the Hugging Face Hub, and the corpus amounts to approximately 70.5 billion tokens per epoch.
Explore the Schneider-Kamp Lab dataset collection on Hugging Face.