Speech-to-Text API that gets every word right
It's simple. We wanted to build the most accurate and fastest transcription software and we delivered on our promise. Find more data below.
Rated 4.9/5 by our users
Trusted by hundreds of fast-growing companies
How it works
How the Speech-to-Text API works
- 01

Send the audio
POST a file or a link to the batch endpoint, or open a WebSocket for live audio. Diarization, sentiment, summaries and your own prompts are switches on the same request — not separate services to stitch together.
- 02

We transcribe it
About a minute of processing per hour of audio, at 95%+ accuracy in 50+ languages. The language is detected for you, and code-switching mid-sentence is handled without configuration.
- 03

Read the JSON
Every word comes back with its timing, confidence, language and speaker. Poll for the result or take a webhook, then export to TXT, DOCX, PDF, SRT or JSON.
Independent benchmark
Independently benchmarked against Microsoft, Google and Whisper
Researchers at the University of Bucharest scored seven transcription systems on over 126 hours of Romanian speech. Vatis finished top two of the seven overall — first on Antena 1 news and first on podcasts, and more accurate than Google on every single set.
- ~95%
Accurate on Antena 1 news
95.6% on the Antena 1 news set — the highest of all seven systems tested, ahead of Microsoft, Google and every Whisper model.
- ~90%
Accurate on podcasts
89.8% on podcasts — first again on unscripted conversation, and more than 12 points clearer than Google on the same audio.
- 6 of 6
Domains ahead of Google
Vatis transcribed more accurately than Google's Chirp on every set in the study, from studio news to overlapping film dialogue.
| Model | In-domain | Out-of-distribution | ||||
|---|---|---|---|---|---|---|
| ProTV — Broadcast news | Antena 1 — Broadcast news | Audiobooks — Read literary prose | Films — Overlapping film dialogue | Stories — Children's stories | Podcasts — Spontaneous conversation | |
| Wav2Vec 2.0 | 27.7 | 25.6 | 40.8 | 75.4 | 54.1 | 36.3 |
| Whisper Small | 31.6 | 26.8 | 40.0 | 60.0 | 41.1 | 31.9 |
| Whisper Large | 12.3 | 5.9 | 14.8 | 27.3 — best in column | 10.9 — best in column | 11.7 |
| Whisper Small + Echo | 9.4 | 10.1 | 18.8 | 54.1 | 21.0 | 21.6 |
| Microsoft Transcribe | 2.9 — best in column | 4.8 | 10.6 — best in column | 31.1 | 17.6 | 11.5 |
| Google Chirp (USM) | 12.1 | 11.3 | 20.2 | 37.6 | 22.4 | 22.4 |
| Vatis | 5.2 | 4.4 — best in column | 13.0 | 31.2 | 16.0 | 10.2 — best in column |
Figures are Word Error Rate (%) — the share of words a system gets wrong, so lower is better. It is the metric the paper reports, kept here so the table can be checked against it line for line; the headline numbers above are its flip side, accuracy. Vatis at 4.4 WER on Antena 1 is 95.6% of words correct. Every score is zero-shot: no system was tuned on these recordings before it was measured.
Outlined values are the best score in their column.
Vatis is a Romanian commercial ASR service known for its strong transcription accuracy in Romanian. Its models are trained on proprietary data and optimized for production use.
Features
Everything the Speech-to-Text API does
Transcription, deployment, languages, customization and transcript metadata. Audio intelligence — summaries, sentiment, topics, intent and PII redaction — has a page of its own.
Transcription: 95%+ Accuracy
Our robust automatic speech recognition (ASR) engine consistently achieves a speech-to-text accuracy exceeding 95%, approaching human transcription on high-quality audio.

Batch Transcription
Accelerate high-volume transcription tasks with our efficient batch transcription API. Process multiple audio and video files simultaneously and receive accurate results in minutes.

Real-Time Transcription
Power real-time workflows with our real-time transcription API. Ideal for live broadcasts, streaming events, and interactive applications.
Deployment

On-Cloud
Simplify deployment with our flexible cloud-based solution. Rapid integration and smooth scalability, perfect for fast-moving teams.

On-Premise
Maintain maximum control with our on-premise deployment option. Ideal for security-sensitive applications and custom integrations.
Languages

Coverage: 50+ languages
Enhance your applications with our transcription services that support over 50 languages. Transcribe content in multiple languages and engage a global audience.

Translation: 30+ languages
Break down language barriers with seamless translation. Convert your transcripts into 30+ languages, boosting accessibility and content reach.

Automatic Language Detection
Eliminate manual language selection — our intelligent API automatically identifies spoken languages.

Real-time Language Switch
Understands more than 50 languages that can be spoken in the same audio input and switches between them in real time as the language changes in the audio.
Customization

Custom Vocabulary
Adapt transcription to your industry with custom vocabulary. Improve accuracy for specialized terminology, jargon, and proper nouns.
Easily add domain-specific terms to our models to ensure that your transcriptions are accurate and relevant. This feature is particularly beneficial for industries like legal, medical, and technical fields where specialized language is common.

Custom Models
Boost transcription accuracy by 10-20%. Fine-tune speech recognition for your unique audio conditions and terminology. Train custom models with your data for unmatched precision.
Our team collaborates with you to create models tailored to your unique needs, ensuring superior performance for niche industries and specialized audio environments.
Train a custom model
Transcript Readability

Numeral Formatting
Ensure clear transcripts with proper numeral formatting. Automatically structure numbers for easy comprehension of dates, currencies, and measurements.

Punctuation and Capitalization
Enhance transcript readability with automatic punctuation and capitalization. Produce professionally formatted text ready for analysis and sharing.

Profanity and Disfluency
Control transcript output with optional profanity filtering and disfluency handling. Create polished results suitable for diverse audiences.

Speaker & Channel Diarization
Identify who said what and when with accurate AI speaker labelling or channel-based labelling. Both batch and real-time transcription.
Transcript Metadata

Word Timestamps
Pinpoint specific moments with word-level timestamps. Quickly navigate audio/video and verify context.

Confidence Scores
Assess transcription accuracy at a glance with confidence scores. Focus editing efforts on sections needing refinement.
API

Multiple Upload Formats
24 audio and video file formats. Conveniently upload common audio and video formats for transcription.

Multiple Export Formats
Easily integrate transcripts into your workflow with flexible export options. Choose the format that best suits your analysis needs: json, txt, pdf, word, srt.

Easy-to-follow Docs
Start fast with our clear and comprehensive API documentation. Quickly implement features and accelerate your development process.
Read the documentation
Case studies
Why teams choose Vatis over everything else
Broadcasting, contact centres, medical, legal, newsrooms, media monitoring, podcasting and research — eight published case studies, each one a team that put Vatis into production.
Languages and formats
Transcribe audio to text in these languages and formats
Supported Languages
English
Spanish
German
French
Italian
Romanian
Portuguese
Polish
Indonesian
Malay
Swedish
Danish
Dutch
Finnish
Norwegian
RussianCatalan
Turkish
Korean
Thai
Japanese
Czech
Ukrainian
Croatian
Greek
Arabic
Bosnian
Hungarian
Bulgarian
Serbian
MacedonianGalician
Cantonese
Slovak
Hindi
Slovenian
Latvian
Urdu
Azerbaijan
Estonian
Vietnamese
Hebrew
Lithuanian
Formats to transcribe audio to text
Formats to transcribe video to text
Export formats for audio and video transcription
Demo
See for yourself

The difference was clear right from the start. Vatis was faster, more accurate, and has only gotten better. It saves us time every day.
I discovered Vatis Tech a year ago, after testing several other speech-to-text solutions. I can honestly say that the difference was noticeable right from the start. Vatis was faster and more accurate than any of the other solutions I tried. A year later, I can say that it has only gotten better. The transcription speed is now even faster, and the accuracy is even higher. Sometimes it surprises me how well Vatis understands, even if the sound quality isn't the best.
It's the perfect solution for our needs and it has saved us so much time and hassle. I highly recommend Vatis Tech to anyone who needs a reliable and accurate speech-to-text solution.
For engineers who read the docs before the marketing page
Read the documentation, try it free, tell us how it goes.
Frequently asked questions
Can't find the answer you're looking for? Reach out to our support team.
What makes Vatis different from Deepgram, AssemblyAI, or Google Speech-to-Text?
Three things. First, real-time multilingual code-switching: our model automatically detects and switches between languages mid-conversation without configuration. Most competitors require you to pre-select a language. Second, built-in audio intelligence (sentiment, topics, intent, PII redaction) in a single API call, no separate services to stitch together. Third, true on-premise deployment for organizations that can't send data to the cloud.
How accurate is the Vatis speech-to-text API?
95%+ on clean audio across all supported languages. Custom vocabulary and custom models can improve accuracy by 10-20% for specialized domains.
Is there a free tier?
Yes. The transcription software opens with 10 free minutes, and the API opens with €10 of credit. Contact us if you need more for testing and we will set you up. The free tier includes all features: transcription, diarization, sentiment analysis, audio intelligence, real-time streaming, and all 50+ languages. No feature gating.
Can I deploy on-premise?
Yes. Vatis offers full on-premise deployment: the entire speech engine runs on your hardware and zero data leaves your network. We also offer private cloud deployment in your AWS, GCP, or Azure environment. That makes Vatis one of very few speech-to-text providers with cloud, private cloud and on-premise options.
What languages are supported for transcription?
Vatis Tech supports transcription in 50+ languages including English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Arabic, Japanese, Korean, Chinese, Hindi, Turkish, Polish, Romanian, Swedish, Danish, Norwegian, Finnish, Czech, Greek, Hungarian, Indonesian, Thai, Vietnamese, Hebrew, and many more. You can also translate transcripts into 30+ languages with one click.
How does real-time streaming work?
Open a WebSocket connection to our streaming endpoint. Send audio chunks (PCM, WAV, or OGG). Receive partial and final transcript events in real time, at roughly 700 milliseconds of latency. Speaker diarization and language detection work in streaming mode. See our streaming quickstart guide for code examples.
Is it secure enough for healthcare and legal applications?
Yes. ISO 27001 certified. GDPR and LGPD compliant. SOC 2 Type II in progress. End-to-end encryption. On-premise deployment ensures PHI and PII never leave your infrastructure. Custom BAA agreements available for HIPAA-covered entities.
What audio formats are supported?
24 formats: MP3, WAV, M4A, FLAC, AAC, OGG, AIFF and WMA for audio; MP4, MKV, AVI, MOV, WebM, WMV, FLV and MPEG for video. Files up to 5GB and 10 hours. Batch processing supports thousands of concurrent files.
What is a Speech-to-Text API?
A speech-to-text API converts spoken language from audio or video files into written text via a programmable interface. Developers integrate it into applications, products, and workflows. Vatis Tech's speech-to-text API goes beyond basic transcription: it includes speaker diarization, sentiment analysis, topic detection, PII redaction, and real-time streaming across 50+ languages.
Which SDKs and integrations are available?
The API is plain REST, so it works from any language that can make an HTTP request, and we publish Python and JavaScript SDKs on top of it. Webhooks notify your application when a batch transcription finishes, and the streaming endpoint is a standard WebSocket. Everything is documented with runnable examples in our API documentation.













