Auto captions in Bengali and Hindi: what actually works

Hundreds of millions of people speak Bengali and Hindi, yet auto captions in these languages are often unusable — English words, Latin spellings or the same phrase repeated. While building VDOAPP's caption tool we dug into why, and what fixes it.

1 October 20267 min read

The symptoms

When we first ran Bengali news audio and Hindi speech recordings through the smaller Whisper models, the results fell into three familiar failure patterns:

  1. Romanised text. Bengali written in Latin letters — the way people type in chat apps — instead of Bengali script.
  2. Silent translation. The model "helpfully" produced English, even though we'd asked for a transcript.
  3. Loops. One phrase repeated line after line, especially over pauses or background music.

Any of these makes captions useless for viewers who read the language — and many creators assume that's just how AI captions are for their language.

Why it happens

Whisper learned from audio paired with text found on the web. English dominates that data, followed by major European and East Asian languages. For Bengali and Hindi, the model saw far less, and much of the text it saw was informal — romanised, mixed with English, or subtitles translated into English.

A small model has limited capacity, so it leans on the patterns it saw most often. When unsure, the "safest" continuation is Latin letters or English. Larger models have room to learn the rarer patterns properly: the right script, the right vocabulary, fewer loops.

We also found an interesting trap. A common trick for stopping repetition — forbidding the model from repeating short word sequences — works in English but damages Indic scripts, where characters combine in ways that look like repetition to the algorithm. It had to go.

What fixed it

  • Using the largest model where it matters. With Whisper large-v3-turbo (our "Max" setting, which runs on WebGPU), Bengali came out in proper Bengali script with only occasional spelling slips. VDOAPP now switches to Max automatically for languages outside the set small models handle well — or to "Best" when a device lacks WebGPU — and says so on screen.
  • Real language detection. Instead of assuming English, "Detect" asks the model which language token is most likely for the opening speech, and uses that for the whole clip.
  • Speech-only windows. Captioning only the stretches that contain voice, and discarding lines that sit over silence, removed most invented text and loops.
  • Script checks. Lines are checked for the expected script, and runaway repeated phrases are collapsed.

Tips for creators

  1. Choose the language yourself for short clips, rather than relying on detection.
  2. Use Chrome or Edge on a computer for the Max model; it needs WebGPU. It downloads once (about 733 MB) and is then cached.
  3. Clean the audio first. Background noise pushes any model towards guessing. VDOAPP's noise removal takes seconds.
  4. Keep music low under speech, or caption only the voice track using "Caption these sounds".
  5. Proofread names and English loanwords. Code-switching (mixing Hindi and English mid-sentence) is still the hardest case for every model.
  6. Want English subtitles too? Choose English under "Show captions in" and Whisper translates as it listens.

The takeaway

Auto captions in Bengali and Hindi aren't inherently bad — small models are. With the right model and a little care over audio, AI captions become a quick proofread rather than a rewrite. And because VDOAPP runs the model on your own device, you can caption as much as you like without paying per minute.

For the full walkthrough, see accurate auto captions in any language.