OdenseNLP logo OdenseNLP
  • Home
  • People
  • Publications
  • Models
  • Data
  • Projects
  • Repositories
  • News
  • Contacts
  • Students

Data

Datasets

Danish, English, multilingual Active collection

Post-Training Dataset Collection

An extensive collection for language-model training with objectives spanning English and Danish instruction and knowledge, mathematics, and agentic-style post-training. The collection contains 86 datasets, with individual datasets reaching up to 2.5 GB of compressed data.

View full collection ↗
Danish Active

Danish Dynaword

A continually expanded corpus of openly licensed Danish free-form text from diverse domains. Version 1.2.20 contains 7.25 million documents and 6.99 billion Llama 3 tokens across 50 dataset subsets.

View dataset ↗

Benchmarks

Danish Active

DaLA

Danish linguistic acceptability benchmark with corrupted and non-corrupted sentences, published as ~8.68k examples with train/validation/test and full-training splits.

View benchmark ↗
Danish Active

SDU-Daisy

Danish-culture benchmark based on the Danish Culture Canon, with 746 closed question-answer pairs for evaluating LLM cultural understanding.

View benchmark ↗
Danish Active

GEC DaLA

Grammatical error correction version of DaLA, pairing original and corrupted Danish sentences with corruption types and affected-token annotations. It provides train, validation, and test splits containing 1,664 examples in total.

View benchmark ↗

© 2026 OdenseNLP · University of Southern Denmark (SDU)

Contact via: petersk@imada.sdu.dk