poly
Research
Benchmarks
About
Collaborate
Research
Benchmarks
About
Collaborate
Research notes
(03)
2026 · 09
MT metrics: what BLEU, chrF and COMET do not measure
A single major accuracy error costs 6.5 chrF; a minor style error costs 2.1. Measuring what MT metrics charge for changing what a text means.
machine translation · metrics
→
2026 · 08
What word error rate does not measure
Word error rate counts every mistake equally and averages over every speaker — a review of what the field's default metric leaves unmeasured.
speech recognition · metrics
→
2026 · 07
Fluent, confident, and wrong: the silent failure in machine translation
How low-resource languages get silently absorbed into better-resourced siblings — and how poly engineers against it.
low-resource languages · safety
→