Translation Quality Assurance for Multilingual LLM Evaluation
Diagnosing translated benchmarks, and quantifying how translation errors affect multilingual LLM evaluation.
Diagnosing translated benchmarks, and quantifying how translation errors affect multilingual LLM evaluation.
Open, multilingual LLMs for Europe — and the tokenizer, instruction-tuning, and evaluation work behind them. Four publications from the OpenGPT-X project (2022–2025).