ACL 2026 preview: Translation errors and multilingual LLM evaluation

The ACL main conference starts soon, so this post is still a short preview. I will add conference impressions and personal takeaways after the presentation. For now, this page collects the core idea of our paper, the public artefact repository, and the citation. ...

July 3, 2026

Translation Quality Assurance for Multilingual LLM Evaluation

Diagnosing translated benchmarks, and quantifying how translation errors affect multilingual LLM evaluation.

July 3, 2026

LREC 2026: Diagnosing translated benchmarks

In May 2026, I presented our work “Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite” at LREC in Palma, Mallorca. The main theme I took away from the conference was that multilingual evaluation is moving beyond simply translating English benchmarks. We also need to ask whether the translated evaluation data is structurally sound, semantically reliable, and documented well enough to support fair model comparisons. This post briefly summarizes our paper and my main personal takeaways from the conference. ...

May 15, 2026

Teuken-7B and multilingual evaluation

Open, multilingual LLMs for Europe — and the tokenizer, instruction-tuning, and evaluation work behind them. Four publications from the OpenGPT-X project (2022–2025).

January 1, 2026