Extensive open training-data collection on Hugging Face
We published an extensive mix of data with different objectives, ranging from English and Danish instruction and knowledge to mathematics and agentic-style post-training data.
OdenseNLP
Safe, Efficient and Open Natural Language Processing @ University of Southern Denmark
We published an extensive mix of data with different objectives, ranging from English and Danish instruction and knowledge to mathematics and agentic-style post-training data.
The Allen Institute for AI (Ai2) recently highlighted FlexMoRE, a new approach to building more efficient modular language models developed by researchers at OdenseNLP and collaborat...
OdenseNLP was at LREC 2026 and authored/contributed in three papers:
Assistant Professor Lukas Galke Poech from OdenseNLP is leading MIST: Scalable Mechanistic Interpretability for Safe and Trustworthy LLM Agents, a project focused on making language ...
Language technologies for low-resource langauges, particularly Danish and neighboring Scandinavian languages.
Fast and efficient NLP architectures and methods.
Making AI systems more safe, trustworthy, and interpretable.
Pipeline for building and evaluating psychologically informed refusal behavior in large language models through data creation, prompting, fine-tuni...
Framework for evaluating memorization and propensity-aware memorization of training data in large language models.
Source code for creating the Danish Corpus of Linguistic Acceptability, designed to evaluate Danish linguistic acceptability with real-world errors.
An extensive collection for language-model training with objectives spanning English and Danish instruction and knowledge, mathematics, and agentic...
A continually expanded corpus of openly licensed Danish free-form text from diverse domains. Version 1.2.20 contains 7.25 million documents and 6.9...
Danish linguistic acceptability benchmark with corrupted and non-corrupted sentences, published as ~8.68k examples with train/validation/test and f...
Danish-culture benchmark based on the Danish Culture Canon, with 746 closed question-answer pairs for evaluating LLM cultural understanding.
Grammatical error correction version of DaLA, pairing original and corrupted Danish sentences with corruption types and affected-token annotations....