I am a Machine Learning Engineer / NLP researcher at TU Dresden, working on multilingual LLM evaluation and scalable benchmarking pipelines. My background is in computer science (university degree) with a background in data management (Big Data architectures & analytics), now focusing on NLP and large language models (LLMs). My research focuses on the reliable evaluation of multilingual LLMs, especially the validity of translated benchmarks, translation-aware evaluation, and culturally robust multilingual benchmarking.
Experience
- TU Dresden (ZIH/VDR) โ Machine Learning Engineer (NLP/LLMs), since 01/2023
- TU Dresden (ScaDS.AI) โ Research Associate (NLP), 02/2020โ02/2021
- Fraunhofer IAIS โ Data Scientist / Lecturer (Big Data architectures), 01/2017โ06/2019
- Fraunhofer IAIS โ Software Engineer / Big Data Architect, 07/2013โ12/2016
- University of Bonn (EIS) โ Research Associate (Semantic Web), 01/2014โ11/2015
Research Interests
- Multilingual LLM evaluation & benchmark design
- Translation artifacts & translation-aware metrics
- Automated benchmark QA (e.g., MT quality estimation, LLM-as-a-judge)
- Human-aligned evaluation / preference validation
- Efficient large-scale evaluation on HPC clusters
Contact: klaudia-doris.thellmann [at] tu-dresden [dot] de
News
- 2026 ACL 2026 โ Quantifying the impact of translation errors on multilingual LLM evaluation
- 2026 LREC 2026 โ Diagnosing translated benchmarks using automated TQE approaches
- 2025 EACL 2025 โ Teuken-7B-Base & Teuken-7B-Instruct: European LLMs
- 2024 EMNLP 2024 โ Investigating multilingual instruction-tuning
- 2024 NAACL 2024 โ Tokenizer choice for LLM training