Speech-to-Text API that gets every word right

It's simple. We wanted to build the most accurate and fastest transcription software and we delivered on our promise. Find more data below.

Rated 4.9/5 by our users

Trusted by hundreds of fast-growing companies

  • ENGIE
  • Groupama
  • Distrigaz Sud Rețele
  • CGS
  • SmartBill
  • Mediatel Data
  • Click Phone

How it works

How the Speech-to-Text API works

One request in, one JSON response out.
  • 01The Vatis Tech API dashboard, with an audio file staged for transcription and the diarization, sentiment, summary and custom-prompt options beside it

    Send the audio

    POST a file or a link to the batch endpoint, or open a WebSocket for live audio. Diarization, sentiment, summaries and your own prompts are switches on the same request — not separate services to stitch together.

  • 02A finished transcription in the API dashboard, split into timestamped turns labelled by speaker, above the playback controls

    We transcribe it

    About a minute of processing per hour of audio, at 95%+ accuracy in 50+ languages. The language is detected for you, and code-switching mid-sentence is handled without configuration.

  • 03The JSON response for a transcription, each word carrying its start and end time, a confidence score, its language and its speaker

    Read the JSON

    Every word comes back with its timing, confidence, language and speaker. Poll for the result or take a webhook, then export to TXT, DOCX, PDF, SRT or JSON.

Independent benchmark

Independently benchmarked against Microsoft, Google and Whisper

Researchers at the University of Bucharest scored seven transcription systems on over 126 hours of Romanian speech. Vatis finished top two of the seven overall — first on Antena 1 news and first on podcasts, and more accurate than Google on every single set.

  • ~95%

    Accurate on Antena 1 news

    95.6% on the Antena 1 news set — the highest of all seven systems tested, ahead of Microsoft, Google and every Whisper model.

  • ~90%

    Accurate on podcasts

    89.8% on podcasts — first again on unscripted conversation, and more than 12 points clearer than Google on the same audio.

  • 6 of 6

    Domains ahead of Google

    Vatis transcribed more accurately than Google's Chirp on every set in the study, from studio news to overlapping film dialogue.

Zero-shot Word Error Rate as a percentage on the in-domain and out-of-distribution test sets. Lower is better.
ModelIn-domainOut-of-distribution
ProTVBroadcast newsAntena 1Broadcast newsAudiobooksRead literary proseFilmsOverlapping film dialogueStoriesChildren's storiesPodcastsSpontaneous conversation
Wav2Vec 2.027.725.640.875.454.136.3
Whisper Small31.626.840.060.041.131.9
Whisper Large12.35.914.827.3 — best in column10.9 — best in column11.7
Whisper Small + Echo9.410.118.854.121.021.6
Microsoft Transcribe2.9 — best in column4.810.6 — best in column31.117.611.5
Google Chirp (USM)12.111.320.237.622.422.4
Vatis5.24.4 — best in column13.031.216.010.2 — best in column

Figures are Word Error Rate (%) — the share of words a system gets wrong, so lower is better. It is the metric the paper reports, kept here so the table can be checked against it line for line; the headline numbers above are its flip side, accuracy. Vatis at 4.4 WER on Antena 1 is 95.6% of words correct. Every score is zero-shot: no system was tuned on these recordings before it was measured.

Outlined values are the best score in their column.

Vatis is a Romanian commercial ASR service known for its strong transcription accuracy in Romanian. Its models are trained on proprietary data and optimized for production use.
Diaconu, Vînaga & Alexe · University of Bucharest · arXiv:2603.02368

Features

Everything the Speech-to-Text API does

Transcription, deployment, languages, customization and transcript metadata. Audio intelligence — summaries, sentiment, topics, intent and PII redaction — has a page of its own.

Transcription: 95%+ Accuracy

Our robust automatic speech recognition (ASR) engine consistently achieves a speech-to-text accuracy exceeding 95%, approaching human transcription on high-quality audio.

  • Batch Transcription

    Accelerate high-volume transcription tasks with our efficient batch transcription API. Process multiple audio and video files simultaneously and receive accurate results in minutes.

  • Real-Time Transcription

    Power real-time workflows with our real-time transcription API. Ideal for live broadcasts, streaming events, and interactive applications.

Deployment

  • On-Cloud

    Simplify deployment with our flexible cloud-based solution. Rapid integration and smooth scalability, perfect for fast-moving teams.

  • On-Premise

    Maintain maximum control with our on-premise deployment option. Ideal for security-sensitive applications and custom integrations.

Languages

  • Coverage: 50+ languages

    Enhance your applications with our transcription services that support over 50 languages. Transcribe content in multiple languages and engage a global audience.

  • Translation: 30+ languages

    Break down language barriers with seamless translation. Convert your transcripts into 30+ languages, boosting accessibility and content reach.

  • Automatic Language Detection

    Eliminate manual language selection — our intelligent API automatically identifies spoken languages.

  • Real-time Language Switch

    Understands more than 50 languages that can be spoken in the same audio input and switches between them in real time as the language changes in the audio.

Customization

  • Custom Vocabulary

    Adapt transcription to your industry with custom vocabulary. Improve accuracy for specialized terminology, jargon, and proper nouns.

    Easily add domain-specific terms to our models to ensure that your transcriptions are accurate and relevant. This feature is particularly beneficial for industries like legal, medical, and technical fields where specialized language is common.

  • Custom Models

    Boost transcription accuracy by 10-20%. Fine-tune speech recognition for your unique audio conditions and terminology. Train custom models with your data for unmatched precision.

    Our team collaborates with you to create models tailored to your unique needs, ensuring superior performance for niche industries and specialized audio environments.

    Train a custom model

Transcript Readability

  • Numeral Formatting

    Ensure clear transcripts with proper numeral formatting. Automatically structure numbers for easy comprehension of dates, currencies, and measurements.

  • Punctuation and Capitalization

    Enhance transcript readability with automatic punctuation and capitalization. Produce professionally formatted text ready for analysis and sharing.

  • Profanity and Disfluency

    Control transcript output with optional profanity filtering and disfluency handling. Create polished results suitable for diverse audiences.

  • Speaker & Channel Diarization

    Identify who said what and when with accurate AI speaker labelling or channel-based labelling. Both batch and real-time transcription.

Transcript Metadata

  • Word Timestamps

    Pinpoint specific moments with word-level timestamps. Quickly navigate audio/video and verify context.

  • Confidence Scores

    Assess transcription accuracy at a glance with confidence scores. Focus editing efforts on sections needing refinement.

API

  • Multiple Upload Formats

    24 audio and video file formats. Conveniently upload common audio and video formats for transcription.

  • Multiple Export Formats

    Easily integrate transcripts into your workflow with flexible export options. Choose the format that best suits your analysis needs: json, txt, pdf, word, srt.

  • Easy-to-follow Docs

    Start fast with our clear and comprehensive API documentation. Quickly implement features and accelerate your development process.

    Read the documentation

Demo

See for yourself

We've prepared a short video tutorial to showcase the seamless functionality of our speech analytics software
AGERPRES

The difference was clear right from the start. Vatis was faster, more accurate, and has only gotten better. It saves us time every day.

I discovered Vatis Tech a year ago, after testing several other speech-to-text solutions. I can honestly say that the difference was noticeable right from the start. Vatis was faster and more accurate than any of the other solutions I tried. A year later, I can say that it has only gotten better. The transcription speed is now even faster, and the accuracy is even higher. Sometimes it surprises me how well Vatis understands, even if the sound quality isn't the best.

It's the perfect solution for our needs and it has saved us so much time and hassle. I highly recommend Vatis Tech to anyone who needs a reliable and accurate speech-to-text solution.

Veronica TudorDeputy Chief Editor, AGERPRES

For engineers who read the docs before the marketing page

Read the documentation, try it free, tell us how it goes.

Frequently asked questions

Can't find the answer you're looking for? Reach out to our support team.

What makes Vatis different from Deepgram, AssemblyAI, or Google Speech-to-Text?

Three things. First, real-time multilingual code-switching: our model automatically detects and switches between languages mid-conversation without configuration. Most competitors require you to pre-select a language. Second, built-in audio intelligence (sentiment, topics, intent, PII redaction) in a single API call, no separate services to stitch together. Third, true on-premise deployment for organizations that can't send data to the cloud.

How accurate is the Vatis speech-to-text API?

95%+ on clean audio across all supported languages. Custom vocabulary and custom models can improve accuracy by 10-20% for specialized domains.

Is there a free tier?

Yes. The transcription software opens with 10 free minutes, and the API opens with €10 of credit. Contact us if you need more for testing and we will set you up. The free tier includes all features: transcription, diarization, sentiment analysis, audio intelligence, real-time streaming, and all 50+ languages. No feature gating.

Can I deploy on-premise?

Yes. Vatis offers full on-premise deployment: the entire speech engine runs on your hardware and zero data leaves your network. We also offer private cloud deployment in your AWS, GCP, or Azure environment. That makes Vatis one of very few speech-to-text providers with cloud, private cloud and on-premise options.

What languages are supported for transcription?

Vatis Tech supports transcription in 50+ languages including English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Arabic, Japanese, Korean, Chinese, Hindi, Turkish, Polish, Romanian, Swedish, Danish, Norwegian, Finnish, Czech, Greek, Hungarian, Indonesian, Thai, Vietnamese, Hebrew, and many more. You can also translate transcripts into 30+ languages with one click.

How does real-time streaming work?

Open a WebSocket connection to our streaming endpoint. Send audio chunks (PCM, WAV, or OGG). Receive partial and final transcript events in real time, at roughly 700 milliseconds of latency. Speaker diarization and language detection work in streaming mode. See our streaming quickstart guide for code examples.

Is it secure enough for healthcare and legal applications?

Yes. ISO 27001 certified. GDPR and LGPD compliant. SOC 2 Type II in progress. End-to-end encryption. On-premise deployment ensures PHI and PII never leave your infrastructure. Custom BAA agreements available for HIPAA-covered entities.

What audio formats are supported?

24 formats: MP3, WAV, M4A, FLAC, AAC, OGG, AIFF and WMA for audio; MP4, MKV, AVI, MOV, WebM, WMV, FLV and MPEG for video. Files up to 5GB and 10 hours. Batch processing supports thousands of concurrent files.

What is a Speech-to-Text API?

A speech-to-text API converts spoken language from audio or video files into written text via a programmable interface. Developers integrate it into applications, products, and workflows. Vatis Tech's speech-to-text API goes beyond basic transcription: it includes speaker diarization, sentiment analysis, topic detection, PII redaction, and real-time streaming across 50+ languages.

Which SDKs and integrations are available?

The API is plain REST, so it works from any language that can make an HTTP request, and we publish Python and JavaScript SDKs on top of it. Webhooks notify your application when a batch transcription finishes, and the streaming endpoint is a standard WebSocket. Everything is documented with runnable examples in our API documentation.