How on-device AI video editing works

Most AI video tools send your footage to a server. VDOAPP does the opposite: it sends the AI to your footage. This guide explains how that's possible in a web page, what gets downloaded, what never leaves your device, and the trade-offs involved.

Updated 1 October 202611 min readAll levels

In short: VDOAPP downloads open AI models (once) from public hosts and runs them on your own processor and graphics card using WebGPU and WebAssembly. Your videos, photos and audio are never uploaded. The costs: a one-time download per model and speed that depends on your device.

Cloud AI vs on-device AI

Cloud AIOn-device AI (VDOAPP)
Your footageUploaded to a company's serversStays on your device
WaitingUpload time + queue + downloadOne-time model download, then processing only
Cost to runServer GPUs — usually paid through subscriptions or creditsYour own hardware — so it can be free
Works offlineNoYes, once models are cached
Model sizeCan be enormousMust fit in a browser tab

How a browser runs a neural network

A neural network is, at heart, a huge list of numbers (weights) and a recipe for multiplying your data by them. Browsers now have two fast ways to do that maths:

  • WebGPU gives web pages access to the graphics card, which can do thousands of multiplications in parallel. It's available in current Chrome and Edge and increasingly in other browsers. With WebGPU, large models like Whisper large-v3-turbo become practical.
  • WebAssembly runs compiled code on the processor at close to native speed. It's the fallback when WebGPU isn't available, and it's what small, frequently-run models (noise suppression, face detection) often use.

Heavy models run in a Web Worker — a background thread — so the editor stays responsive while, say, captions are generated. Libraries such as Transformers.js (for Whisper, translation, depth and background removal), MediaPipe (for people, faces and face landmarks) and kokoro-js (for speech) turn the model files into something the browser can execute.

The models inside VDOAPP

FeatureModelApprox. download
Auto captionsOpenAI Whisper (tiny, base, small, large-v3-turbo)39 MB – 733 MB depending on size and device
Caption translationOpus-MT (Helsinki-NLP), one small model per languageTens of MB each
Text to speechKokoro 82M88 MB
People cut-out, reframeMediaPipe selfie segmenterUnder 1 MB
Face blur, face stickers, thumbnailsMediaPipe face detector and face landmarkerA few MB
3D photo, portrait blurDepth Anything V2 small26 MB
Any-subject photo cut-outBRIA RMBG-1.442 MB
Noise removalRNNoiseSmall, bundled in its library

Some features use no neural network at all: silence detection, beat detection, scene-change detection and stabilisation are classic signal processing written in plain JavaScript — fast, predictable and tiny.

Downloads and caching

Models are fetched only when you first use the feature that needs them, from public model hosts (Hugging Face, Google's storage for MediaPipe models, and the jsDelivr CDN for libraries). Your browser stores them in its cache, so the next use starts immediately and works offline. Clearing your browser's site data removes them; they'll simply download again when needed.

Downloading a model sends an ordinary web request to that host — like loading an image — so the host sees your IP address, as with any website. Nothing about your video is included in that request.

What affects speed

  • WebGPU or not. The biggest factor for captions: with a recent graphics chip, even the Max model runs faster than real time.
  • Memory. Large models need a few gigabytes of RAM. On phones, stick to smaller models.
  • First use. The first run includes the download and compiling the model for your hardware.
  • Other tabs. Close heavy tabs when captioning long videos.

What stays private

  • Your video, photo and audio files — read from your disk by the browser, never uploaded.
  • Everything derived from them — transcripts, face positions, depth maps, masks.
  • Your project — autosaved in your browser's own storage on your device.

The website itself uses analytics and advertising cookies as described in our privacy policy; those relate to your visit, not to your media. The editor page carries no ads.

Honest limitations

  • Browser-sized models are smaller than the largest cloud models, so some tasks — video background removal for arbitrary objects, for instance — aren't yet practical in real time.
  • Performance varies widely between a new laptop and a five-year-old phone.
  • WebGPU support is still arriving in some browsers; without it, big models are slow or unavailable.

Frequently asked questions

How can I check that nothing is uploaded?

Open your browser's developer tools, go to the Network tab and use an AI feature. You'll see model files being downloaded the first time, but no request carrying your media.

Does it work offline?

Once the editor page and the models you use are cached, the AI features work without a connection. The page itself needs to be loaded online first.

Why is the first caption run slow?

It downloads the speech model and prepares it for your hardware. Later runs skip both steps.