In short: VDOAPP downloads open AI models (once) from public hosts and runs them on your own processor and graphics card using WebGPU and WebAssembly. Your videos, photos and audio are never uploaded. The costs: a one-time download per model and speed that depends on your device.
Cloud AI vs on-device AI
| Cloud AI | On-device AI (VDOAPP) | |
|---|---|---|
| Your footage | Uploaded to a company's servers | Stays on your device |
| Waiting | Upload time + queue + download | One-time model download, then processing only |
| Cost to run | Server GPUs — usually paid through subscriptions or credits | Your own hardware — so it can be free |
| Works offline | No | Yes, once models are cached |
| Model size | Can be enormous | Must fit in a browser tab |
How a browser runs a neural network
A neural network is, at heart, a huge list of numbers (weights) and a recipe for multiplying your data by them. Browsers now have two fast ways to do that maths:
- WebGPU gives web pages access to the graphics card, which can do thousands of multiplications in parallel. It's available in current Chrome and Edge and increasingly in other browsers. With WebGPU, large models like Whisper large-v3-turbo become practical.
- WebAssembly runs compiled code on the processor at close to native speed. It's the fallback when WebGPU isn't available, and it's what small, frequently-run models (noise suppression, face detection) often use.
Heavy models run in a Web Worker — a background thread — so the editor stays responsive while, say, captions are generated. Libraries such as Transformers.js (for Whisper, translation, depth and background removal), MediaPipe (for people, faces and face landmarks) and kokoro-js (for speech) turn the model files into something the browser can execute.
The models inside VDOAPP
| Feature | Model | Approx. download |
|---|---|---|
| Auto captions | OpenAI Whisper (tiny, base, small, large-v3-turbo) | 39 MB – 733 MB depending on size and device |
| Caption translation | Opus-MT (Helsinki-NLP), one small model per language | Tens of MB each |
| Text to speech | Kokoro 82M | 88 MB |
| People cut-out, reframe | MediaPipe selfie segmenter | Under 1 MB |
| Face blur, face stickers, thumbnails | MediaPipe face detector and face landmarker | A few MB |
| 3D photo, portrait blur | Depth Anything V2 small | 26 MB |
| Any-subject photo cut-out | BRIA RMBG-1.4 | 42 MB |
| Noise removal | RNNoise | Small, bundled in its library |
Some features use no neural network at all: silence detection, beat detection, scene-change detection and stabilisation are classic signal processing written in plain JavaScript — fast, predictable and tiny.
Downloads and caching
Models are fetched only when you first use the feature that needs them, from public model hosts (Hugging Face, Google's storage for MediaPipe models, and the jsDelivr CDN for libraries). Your browser stores them in its cache, so the next use starts immediately and works offline. Clearing your browser's site data removes them; they'll simply download again when needed.
Downloading a model sends an ordinary web request to that host — like loading an image — so the host sees your IP address, as with any website. Nothing about your video is included in that request.
What affects speed
- WebGPU or not. The biggest factor for captions: with a recent graphics chip, even the Max model runs faster than real time.
- Memory. Large models need a few gigabytes of RAM. On phones, stick to smaller models.
- First use. The first run includes the download and compiling the model for your hardware.
- Other tabs. Close heavy tabs when captioning long videos.
What stays private
- Your video, photo and audio files — read from your disk by the browser, never uploaded.
- Everything derived from them — transcripts, face positions, depth maps, masks.
- Your project — autosaved in your browser's own storage on your device.
The website itself uses analytics and advertising cookies as described in our privacy policy; those relate to your visit, not to your media. The editor page carries no ads.
Honest limitations
- Browser-sized models are smaller than the largest cloud models, so some tasks — video background removal for arbitrary objects, for instance — aren't yet practical in real time.
- Performance varies widely between a new laptop and a five-year-old phone.
- WebGPU support is still arriving in some browsers; without it, big models are slow or unavailable.
Frequently asked questions
How can I check that nothing is uploaded?
Open your browser's developer tools, go to the Network tab and use an AI feature. You'll see model files being downloaded the first time, but no request carrying your media.
Does it work offline?
Once the editor page and the models you use are cached, the AI features work without a connection. The page itself needs to be loaded online first.
Why is the first caption run slow?
It downloads the speech model and prepares it for your hardware. Later runs skip both steps.