One call, four speech engines: why we stayed with Azure
Speech recognition is Sayply's cost of goods: we sell minutes, and every minute is sold to us by an STT provider. So "should we move to a cheaper one" comes up regularly. Instead of going by feel, we measured it.
How we measured
We recorded a real 5 min 23 s call — two channels (microphone and system audio), 16 kHz mono, exactly the format our production pipeline feeds into recognition. We ran both channels through four providers and computed WER — the percentage of words recognized incorrectly against a reference text.
The numbers
| Provider | Your channel (WER) | Other side (WER) |
|---|---|---|
| Azure (our prod) | 15.5% | 24.2% |
| AssemblyAI | 15.5% | 27.8% |
| Deepgram | 20.7% | 35.9% |
| Speechmatics | 29.2% | 31.8% |
The absolute numbers are inflated equally for everyone — the recording deviates from the script in places, and those deviations count as errors. For comparing providers that does not matter: the audio is identical.
What settled it
AssemblyAI matched Azure on quality at 70% lower cost — but does not support Ukrainian in streaming. Its realtime mode covers 18 languages; Ukrainian is not among them. And the workaround — "stream with one provider, re-process the recording with a cheaper one later" — is impossible for us by design: it requires storing audio, which our privacy policy explicitly promises not to do.
Speechmatics garbles numbers. A self-correction — "fifteen… no, eighteen" — came out as "20 seats". For a product that turns conversations into tasks, a wrong number inside a task is the worst class of error there is.
Deepgram is the only real alternative. It holds Ukrainian in streaming, first results in 1.6 s, and costs notably less. But its WER is a third worse, and recognition quality is literally our product.
The decision
We stay on Azure. Switching providers is a lever of scale, not of quality: at current volumes the savings are noise, while recognition gets measurably worse. The benchmark harness is reproducible, so when volumes grow we will re-run it instead of reminiscing.
An honest finding
As a side effect, the benchmark exposed a limitation shared by all four providers: when a channel is pinned to one language, they do not transcribe a foreign-language line verbatim — they translate it. An English "Sounds good, let's say we have five managers…" comes out as its Ukrainian retelling.
For a bilingual call this means: the transcript of foreign-language lines is a retelling, not a verbatim record. Many users actually prefer the retelling in their own language — but we will not call it verbatim. If word-for-word accuracy in both languages matters to you, pick the language per channel: here is how.