Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations

Open Access
Authors
  • S. Tan
  • Ryosuke Mitani
  • Ritvik Choudhary
  • Qiyu Wu
  • Toshiyuki Sekiya
  • C. Monz
Publication date 2026
Host editors
  • M. Liakata
  • V.P. Moreira
  • J. Zhang
  • D. Jurgens
Book title The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026) : Findings of the Association for Computational Linguistics: ACL 2026
Book subtitle Findings 2026 : July 2-7, 2026
ISBN (electronic)
  • 9798891763951
Event 64th Annual Meeting of the Association for Computational Linguistics
Pages (from-to) 7374–7398
Publisher Kerrville, TX: Association for Computational Linguistics
Organisations
  • Faculty of Science (FNWI) - Informatics Institute (IVI)
Abstract
Over the years, automatic MT metrics have hillclimbed benchmarks and presented strong and sometimes human-level agreement with human ratings. Yet they remain black-box, offering little insight into their decision-making and often failing under real-world out-of-distribution (OOD) inputs. We introduce Remedy-R, a reasoning-driven generative MT metric trained with reinforcement learning from pairwise translation preferences, without requiring error-span annotations or distillation from closed LLMs. Remedy-R produces step-by-step analyses of accuracy, fluency, and completeness, followed by a final score, enabling more interpretable assessments. With only 60K training pairs across two language pairs, Remedy-R remains competitive with top scalar metrics and GPT-4-based judges on WMT22-24 meta-evaluation, generalizes to other languages, and exhibits strong robustness on OOD stress tests. Moreover, Remedy-R models generate self-reflective feedback that can be reused for translation improvement. Building on this finding, we introduce Remedy-R Agent, a simple evaluate-revise pipeline that leverages Remedy-R's evaluation analysis to refine translations. This agent consistently improves translation quality across diverse models, including Qwen2.5, ALMA-R, GPT-4o-mini, and Gemini-2.0-Flash, suggesting that Remedy-R's reasoning captures translation-relevant information and is practically useful.
Document type Conference contribution
Language English
Published at
Downloads
2026.findings-acl.364 (Final published version)
Permalink to this page
Back