About
Translation systems fail confidently.
A speech translation can score well on the metrics the field optimizes (WER, BLEU, etc.) and still contain the one mistake that matters: a dosage inverted, a symptom misheard, a testimony subtly changed. After years of working on speech recognition and machine translation, and after conversations with the professionals who live with the consequences of translation errors, we came to a simple conclusion: the errors that matter most are exactly the ones our standard measures don't see.
poly exists to fix that.
source · de
„Nehmen Sie nicht mehr als zwei Tabletten täglich.“
model output · en
“Take more than two tablets daily.”
We build and evaluate speech translation for the settings where errors cost the most, and we don't call a language ready until it is safe there. By making safe translation available in more and more languages, we democratize access to AI — so that everyone, everywhere, can share in its benefits.
Why it's hard
Safe speech translation is hard for reasons that compound each other.
01
Low-resource languages fail silently.
For the languages where translation is needed most, models produce output that is fluent, confident — and wrong. Closely related languages collapse into each other, and the system gives no signal that anything failed. The better the models get at sounding right, the harder these failures are to see.
02
Quality collapses exactly where stakes rise.
General-purpose models are trained on general text. In medical, legal, or administrative settings — precise terminology, high consequences — their precision degrades at the very moment it matters most.
03
Real-time conversation leaves no room to hide.
A conversation cannot wait. Speech recognition, translation, and synthesis must run in seconds, and every optimization that buys speed risks costing accuracy. Achieving both, at the same time, in a live conversation, is an open engineering and research problem.
04
The data doesn't exist.
For low-resource languages there is too little training data, too little evaluation data, and almost none from the real settings where translation errors matter. This is why the field’s blind spot persists: you cannot measure what you have no data to measure with. We address this at the root — through ethical data gathering, building the training and evaluation resources these languages are missing, with the consent and participation of the communities who speak them.
How we work
Evidence over claims. Every model in our pipeline is benchmarked, and swapped when the data says a better one exists. We publish per-language quality ratings that say plainly where translation works and where it struggles — a language ships when it crosses a measured safety bar, not when a roadmap says so.
Verification built in. We don’t ask anyone to trust a black box. Every translation comes with a transcript, so speakers can see what the system heard and catch errors before they matter. For languages that haven’t yet reached the safety bar, that verification surface is mandatory, not optional.
Open-weight, self-hosted. We run every stage of the pipeline ourselves, on our own infrastructure in Switzerland, built entirely on open-weight models. No third-party translation APIs, no data leaving our systems. This keeps our results reproducible, our stack auditable, and the data of the people who rely on us under our control alone.
Data gathered ethically. Where training and evaluation data doesn’t exist, we build it — with the consent and participation of the communities who speak the language, never by scraping them. We retain no conversation data from deployment: what people say through poly is processed, translated, and gone.
Research in the open. We publish our benchmarks, methods, and findings as we go — including negative results. What we claim publicly is what the data shows, with no marketing layer in between.
Founder
poly was founded by Jonathan Gosteli, AI researcher, in Zürich.