Hello ๐Ÿ‘‹

I am a Machine Learning Engineer / NLP researcher at TU Dresden, working on multilingual LLM evaluation and scalable benchmarking pipelines. My background is in computer science (university degree) with a background in data management (Big Data architectures & analytics), now focusing on NLP and large language models (LLMs). My research focuses on the reliable evaluation of multilingual LLMs, especially the validity of translated benchmarks, translation-aware evaluation, and culturally robust multilingual benchmarking.

Experience

  • TU Dresden (ZIH/VDR) โ€” Machine Learning Engineer (NLP/LLMs), since 01/2023
  • TU Dresden (ScaDS.AI) โ€” Research Associate (NLP), 02/2020โ€“02/2021
  • Fraunhofer IAIS โ€” Data Scientist / Lecturer (Big Data architectures), 01/2017โ€“06/2019
  • Fraunhofer IAIS โ€” Software Engineer / Big Data Architect, 07/2013โ€“12/2016
  • University of Bonn (EIS) โ€” Research Associate (Semantic Web), 01/2014โ€“11/2015

Research Interests

  • Multilingual LLM evaluation & benchmark design
  • Translation artifacts & translation-aware metrics
  • Automated benchmark QA (e.g., MT quality estimation, LLM-as-a-judge)
  • Human-aligned evaluation / preference validation
  • Efficient large-scale evaluation on HPC clusters

Contact: klaudia-doris.thellmann [at] tu-dresden [dot] de

Portrait

News

  • 2026 ACL 2026 โ€” Quantifying the impact of translation errors on multilingual LLM evaluation
  • 2026 LREC 2026 โ€” Diagnosing translated benchmarks using automated TQE approaches
  • 2025 EACL 2025 โ€” Teuken-7B-Base & Teuken-7B-Instruct: European LLMs
  • 2024 EMNLP 2024 โ€” Investigating multilingual instruction-tuning
  • 2024 NAACL 2024 โ€” Tokenizer choice for LLM training