Accurate auto captions in any language

Automatic captions are excellent in English and surprisingly poor in many other languages — unless you know how the models work. This guide explains what's going on under the hood and how to get clean, correctly-scripted captions in Hindi, Bengali, Arabic, Spanish or anything else.

Updated 1 October 202610 min readBeginner to intermediate

In short: pick the spoken language instead of relying on detection for short clips; for South Asian, Middle Eastern and African languages use the largest model your device can run; clean the audio first; and budget two minutes to proofread names and numbers.

Why captions matter more than ever

Most short videos are watched with the sound off — on a bus, in bed, in an office. Captions keep those viewers. They also help the roughly one in twenty people with disabling hearing loss, people watching in a second language, and search engines, which can't listen to your video but can read an .srt file.

The trouble is that typing captions by hand takes four to ten times the length of the video. Automatic speech recognition cuts that to a quick proofread — when it works.

How speech models listen

VDOAPP uses Whisper, an open speech-recognition model trained on hundreds of thousands of hours of audio from the web in close to a hundred languages. It works in three steps:

  1. The sound becomes a picture. Audio is resampled to 16,000 samples per second and turned into a spectrogram — a chart of which pitches are loud at each moment.
  2. An encoder "reads" 30 seconds at a time. Whisper always looks at 30-second windows, which is why long videos are processed in pieces.
  3. A decoder writes text, token by token. It first predicts the language, then the words, then timestamps. Each word is a guess conditioned on the previous ones — which is why a model can sometimes "hallucinate" fluent-sounding text that nobody said, especially over silence or music.

VDOAPP guards against the most common failure — words invented over quiet passages — by checking the actual loudness of every caption's time span and discarding lines with no voice under them. It also tightens each caption's start and end to where the voice really begins and stops.

Why some languages go wrong

Whisper's training data is dominated by English, followed by major European and East Asian languages. For a language like Bengali, the model saw a small fraction of the data. The result in the small models is predictable:

  • Romanisation — Bengali or Hindi written in Latin letters ("ami tomake bhalobashi") because that's what much of the web's informal text looks like.
  • Translation instead of transcription — the model drifts into English.
  • Repetition loops — one phrase written again and again.

The fix is model size. The largest model, Max (Whisper large-v3-turbo), has far more capacity for less-represented languages and writes them in their proper script. That's why VDOAPP automatically switches to Max — or Best, if your browser lacks WebGPU — when you choose a language outside the set that the small models handle well, and tells you it has done so.

Real-world test: on a Bengali news clip, the Balanced model produced romanised, partly English text; Max produced correct Bengali script with only a couple of spelling slips.

Choosing a model

Your situationPick
Clear English, one speaker, quick social clipFast or Balanced
Spanish, French, German, Portuguese, Japanese, Chinese…Balanced, Best for accents or noise
Hindi, Bengali, Urdu, Tamil, Telugu, Arabic, Persian, Swahili…Max (needs WebGPU) — VDOAPP selects it for you
Music under the voice, several speakers, jargonBest or Max
Older laptop or phone, long videoBalanced — and caption in parts

Language Detect works by asking the model which language it hears in the first seconds of speech. It's reliable on clear speech but can be fooled by a short clip, a song intro or code-switching (mixing Hindi and English, say). If you know the language, choose it.

Audio beats model size

A bigger model can't recover words that are buried in noise. Before captioning:

  • Remove background noise on noisy clips (AI tab → Voice). Fans and traffic confuse the model and make it guess.
  • Turn down music under speech, or caption a clean voiceover track rather than the final mix — the Caption these sounds list lets you tick only the sources that contain speech.
  • Remove long silences first if you plan to. Captions follow cuts, but captioning the tight edit is quicker.

Step by step in VDOAPP

  1. Finish your cut (or near enough). Add clips, trim, remove silences.
  2. Open the AI tab → Auto captions. Choose the Language and check the Accuracy it selected.
  3. Choose what to caption under Caption these sounds: the clips' sound, voiceovers, or a background track that contains speech (for example a separately recorded interview).
  4. Pick a style and position, then press Generate captions. The first run downloads the model; progress shows in the panel.
  5. Proofread in the Text tab, then export the video — and Export .srt if you'll upload captions separately.

Fixing the last 5%

Even a great model gets some things wrong. Proofread for these, in this order:

  1. Names and brands. The model spells unknown names phonetically.
  2. Numbers. "Fifteen" vs "fifty" is a classic confusion; check prices, dates and statistics.
  3. Homophones and negations. "Can" vs "can't" changes meaning completely and is easy to mishear.
  4. Line breaks. Keep each caption to one idea; split long lines at natural pauses.

Playing the video at normal speed while reading is the fastest way — your ear catches mismatches your eye skims over.

Caption style that reads well

  • Contrast first. White outline suits most footage; Black box is safest over busy or bright backgrounds; Bold yellow grabs attention on social feeds.
  • Mind the interface. On TikTok, Reels and Shorts the bottom fifth of the screen is covered by captions and buttons — use Middle position for vertical videos.
  • Short lines. Two lines at most; around 35–42 characters per line for landscape video, fewer for vertical.
  • Enough time to read. At least about a second per caption, and never faster than people speak.

Frequently asked questions

Why are my Hindi captions in English letters?

Smaller speech models often romanise languages they saw little of during training. Use the Max model (Chrome or Edge with WebGPU) for proper Devanagari script; VDOAPP selects it automatically when you choose Hindi.

How long does captioning take?

After the one-time download, a few minutes of speech usually takes well under a minute on a recent laptop with WebGPU. Without it, larger models are noticeably slower.

Can it handle two languages mixed together?

Partly. Whisper expects one language per 30-second window; mixed speech tends to come out in the dominant language. Choose the main language and correct the other phrases by hand.

Does anything get uploaded?

No. The model runs inside your browser and your audio stays on your device.