Technology
Accuracy that starts at 95%+ and improves on your audio
Vatis transcribes speech it has never heard at 95%+ accuracy. What takes it past that is a training loop we run ourselves — transcribe, correct by hand, re-train, deploy — on the recordings your own team brings.
Independent benchmark
Independently benchmarked against Microsoft, Google and Whisper
Researchers at the University of Bucharest scored seven transcription systems on over 126 hours of Romanian speech. Vatis finished top two of the seven overall — first on Antena 1 news and first on podcasts, and more accurate than Google on every single set.
- ~95%
Accurate on Antena 1 news
95.6% on the Antena 1 news set — the highest of all seven systems tested, ahead of Microsoft, Google and every Whisper model.
- ~90%
Accurate on podcasts
89.8% on podcasts — first again on unscripted conversation, and more than 12 points clearer than Google on the same audio.
- 6 of 6
Domains ahead of Google
Vatis transcribed more accurately than Google's Chirp on every set in the study, from studio news to overlapping film dialogue.
| Model | In-domain | Out-of-distribution | ||||
|---|---|---|---|---|---|---|
| ProTV — Broadcast news | Antena 1 — Broadcast news | Audiobooks — Read literary prose | Films — Overlapping film dialogue | Stories — Children's stories | Podcasts — Spontaneous conversation | |
| Wav2Vec 2.0 | 27.7 | 25.6 | 40.8 | 75.4 | 54.1 | 36.3 |
| Whisper Small | 31.6 | 26.8 | 40.0 | 60.0 | 41.1 | 31.9 |
| Whisper Large | 12.3 | 5.9 | 14.8 | 27.3 — best in column | 10.9 — best in column | 11.7 |
| Whisper Small + Echo | 9.4 | 10.1 | 18.8 | 54.1 | 21.0 | 21.6 |
| Microsoft Transcribe | 2.9 — best in column | 4.8 | 10.6 — best in column | 31.1 | 17.6 | 11.5 |
| Google Chirp (USM) | 12.1 | 11.3 | 20.2 | 37.6 | 22.4 | 22.4 |
| Vatis | 5.2 | 4.4 — best in column | 13.0 | 31.2 | 16.0 | 10.2 — best in column |
Figures are Word Error Rate (%) — the share of words a system gets wrong, so lower is better. It is the metric the paper reports, kept here so the table can be checked against it line for line; the headline numbers above are its flip side, accuracy. Vatis at 4.4 WER on Antena 1 is 95.6% of words correct. Every score is zero-shot: no system was tuned on these recordings before it was measured.
Outlined values are the best score in their column.
Vatis is a Romanian commercial ASR service known for its strong transcription accuracy in Romanian. Its models are trained on proprietary data and optimized for production use.
Training loop
The science behind the magic
- 01
We start with raw data
We transcribe the audio data with the current version of our Speech-to-Text model. We split the result into fragments that can be easily analyzed, corrected, and validated. Also, we start an initial self-supervised training process for our technology at this step.
- 02
We do in-house labelling
Our team of validators starts to analyse, correct and validate the data from the previous step. They take unlabelled data and label it.
- 03
Training, lots of it
When we have enough new hours validated by our team, we re-train the Speech-to-Text model using a supervised technique this time.
- 04
Deployment
When the training has finished, we deploy the new version of the model. We are also constantly researching better ways to improve our model's architecture.
Continuous learning
…then the cycle repeats. Constantly.
Every hour our validators correct becomes training data for the next version, so the model does not stand still between releases. Point the same loop at one customer's own recordings — their vocabulary, their speakers, their recording conditions — and a model tuned to them usually takes one to two months.

Considering that for almost 4 years we have transcribed by hand legal debates and conferences at JURIDICE.ro, Vatis Tech came as a rescue.
We like how easy the platform is to use, the fact that the transcripts are accurate and it takes a few minutes for them to be ready.

More from Vatis
Transcription Software
Learn more about how Vatis Tech helps teams and individuals streamline their transcription workflow.
Learn moreSpeech-to-Text API
Supercharge your apps with Vatis Tech's accurate, accessible and affordable Speech-to-Text API.
Learn moreAudio Intelligence
Extract actionable audio insights in minutes from your transcription API.
Learn more
For engineers who read the docs before the marketing page
Read the documentation, try it free, tell us how it goes.