How to caption a video
- Add your video. Open the editor and drop in your clips. Cut them first if you like — captions are placed on the finished timeline and follow later cuts and speed changes too.
- Choose the language. In the AI tab, under Auto captions, leave Language on Detect or pick the spoken language. For languages the smaller models handle poorly, VDOAPP switches to a larger model automatically and tells you why.
- Pick what to caption. Under Caption these sounds, tick the clips' own sound, a voiceover, or even a track you added as background music — useful when the speech lives in a separate recording.
- Generate and polish. Press Generate captions. Every line becomes a text layer in the Text tab, where you can fix a word, change the timing or restyle it. Export the video, or save the captions as an .srt file.
What makes it different
- Four accuracy levels. Fast, Balanced and Best work on any modern browser; Max (Whisper large-v3-turbo) appears when your browser has WebGPU and gives far better results for languages such as Bengali, Hindi, Urdu and Tamil.
- Real language detection. Detect asks the model which language is being spoken instead of assuming English.
- Translation built in. Show captions in can translate into English, Spanish, French, German, Italian, Hindi, Arabic, Chinese, Japanese, Russian, Indonesian or Vietnamese.
- Tight timing. Each caption is trimmed to where the voice actually starts and stops, and nothing is written over silence — a common problem with speech models, which tend to "hear" words in quiet gaps.
- Three ready styles — white outline, black box and bold yellow — at the bottom, middle or top of the frame.
- .srt in and out. Import subtitles you already have, or export the generated ones for YouTube, Vimeo or a media player.
How big are the models?
Each model downloads once and your browser keeps it, so the second video is much faster than the first. Sizes depend on whether your browser can use the graphics card (WebGPU):
| Accuracy | With WebGPU | Without | Good for |
|---|---|---|---|
| Fast | 114 MB | 39 MB | Clear English, quick drafts |
| Balanced | 197 MB | 73 MB | Most videos in major European and East Asian languages |
| Best | 559 MB | 238 MB | Accents, music underneath, technical words |
| Max | 733 MB | — | South Asian, Middle Eastern and other less-represented languages |
Tip: clean audio beats a bigger model. If there's a fan or traffic behind the voice, run Remove background noise on the clip before captioning.
Frequently asked questions
Is my video uploaded for transcription?
No. The speech model is downloaded to your browser and runs on your own processor or graphics card. The audio of your video never leaves your device.
Which languages are supported?
The language list includes English, Spanish, French, German, Italian, Portuguese, Dutch, Hindi, Urdu, Bengali, Arabic, Turkish, Russian, Chinese, Japanese, Korean, Indonesian, Malay, Persian, Punjabi, Tamil, Telugu, Marathi, Gujarati, Vietnamese, Thai, Polish, Ukrainian, Hebrew, Greek, Swedish, Swahili and Filipino, plus automatic detection.
Why did it choose a bigger model than I picked?
The Fast and Balanced models were trained on little data for some languages and tend to write them in Latin letters or guess words. For those languages VDOAPP uses Max when your browser supports WebGPU, or Best otherwise, and shows a note explaining the switch.
Can I edit the captions?
Yes. Every caption is an ordinary text layer in the Text tab, so you can correct words, change when it appears, or restyle and animate it like any other text.
Does it work offline?
After a model has been downloaded once, captioning with that model works without an internet connection.