The upload problem
Think about what's actually in the footage people edit. Birthday parties with children. The inside of someone's home. Friends who agreed to be filmed for a group chat, not for a company's training data. A dash cam clip that might become evidence. A client's unreleased product.
Cloud AI editors ask you to send all of that to their servers before they'll caption it, cut it out or clean up the sound. Most are careful and well-intentioned. But once a file leaves your device, you're trusting someone else's security, retention policy and business model — and those can change after you've clicked upload.
There's a practical cost too. A ten-minute 4K clip can be several gigabytes. On an ordinary home connection, uploading it takes longer than the AI takes to process it. Then you wait in a queue, and then you download the result.
What changed
For years there was a good reason for the upload: AI models needed server-grade graphics cards. Three things changed that:
- Open models got good and got small. Speech recognition (Whisper), depth estimation (Depth Anything), background removal and natural text-to-speech (Kokoro) are now openly available in sizes that fit in a browser.
- Browsers got access to the graphics card. WebGPU lets a web page run neural networks on the same GPU that plays your games. WebAssembly runs the rest on the processor at close to native speed.
- Browsers learned to handle video directly. WebCodecs gives web pages frame-level access to the device's video decoders and encoders — which is how VDOAPP exports frame-perfect MP4 faster than real time.
Put together, a browser tab can now do what needed a render server a few years ago. So we built the editor around that: the AI comes to your footage, not the other way round.
What it cost us
We won't pretend it was free. Running everything locally meant accepting constraints a cloud product doesn't have:
- Model downloads. The first time you use a tool, its model has to download — up to 733 MB for the most accurate caption model. We show the size up front, cache it, and never download it twice.
- Your device sets the speed. A recent laptop with WebGPU captions a few minutes of speech in well under a minute. An old phone will be much slower. We pick sensible defaults per device, but we can't rent you a faster GPU.
- Language is hard. Small speech models romanise Bengali and Hindi or drift into English. We had to detect that, switch to larger models automatically, and explain the switch — rather than quietly giving poor results.
- Memory limits. Browsers cap how much memory a tab can use, which is why we built preview copies for 4K and process long audio in pieces.
We also gave up a common business model. With no servers doing the heavy lifting, there's no per-minute cost to recover through credits or subscriptions. Advertising on our information pages covers our running costs; the editor itself stays ad-free.
What it means for you
- Privacy you can verify. Open your browser's network panel while captioning: you'll see model files arriving the first time, and nothing carrying your media leaving.
- No queues, no credits. Caption as many videos as you like. The only limit is your patience and your battery.
- Works offline. Once the editor and the models you use are cached, you can edit on a plane.
- Sensitive footage becomes editable. Journalists, teachers and anyone handling footage of other people can use AI tools without a data-processing agreement.
What's next
Browser AI is improving fast. WebGPU support is spreading across browsers, models are getting smaller for the same quality, and devices keep adding dedicated AI hardware. Every improvement there makes VDOAPP faster without us touching your data.
If there's a tool you'd like to see — or a language that doesn't caption well yet — tell us. And if you want the technical detail behind all this, read how on-device AI video editing works.