In short: clean the sound → remove silences → fix the order and hook → add auto zoom and B-roll → caption → grade and export. Each step makes the next one easier.
Why the order matters
Each tool in this workflow depends on the one before. Noise removal makes silence detection accurate. Silence removal makes the video shorter, so captioning is faster and captions are placed on the final timing. Captions give you a transcript to judge structure. Doing it in a different order means redoing work.
1. Fix the sound
Select each clip and use 🧹 Remove background noise (AI tab → Voice) if there's any hum, fan or traffic. Then check levels: if you recorded in several sessions, even out the Volume of each clip in the Clip tab so the voice doesn't jump between cuts. More on voice audio →
2. Set the pace
Use Remove silences in the AI tab. The right pause length depends on your content:
| Style | Pauses longer than | Keep around speech |
|---|---|---|
| Calm lesson, meditation, storytelling | 1.0 s | 0.25 s |
| Typical YouTube explainer | 0.5–0.6 s | 0.15 s |
| Fast-paced short | 0.3–0.4 s | 0.1 s |
Then watch it through once and remove the bad takes by hand: repeated sentences, false starts, tangents. Split with S and delete. When you restart a sentence on camera, the last take is usually the best one.
3. Structure and hook
People decide within seconds whether to stay. Look at your first 15 seconds: does it say what the viewer gets? Common fixes:
- Cut the greeting. "Hey guys, welcome back" can come after the hook, or not at all.
- Move the best line to the start. A surprising result or bold claim from minute three often makes the perfect opening.
- Mark chapters with markers (M) as you watch — they help you see the structure and plan titles.
4. Keep it visual
A static face for five minutes is hard to watch, but you don't need a second camera:
- Auto zoom. 🔍 Auto zoom in the AI tab listens for emphasised moments — where your voice rises in energy — and punches in on them. It turns some jump cuts into deliberate framing changes, like a two-camera shoot.
- B-roll with PiP. In the PiP tab, add a clip or screenshot of what you're talking about. Make it full frame (⛶ Fill frame) to cut away, or keep it small beside you.
- On-screen text for key numbers and terms — short, big, and on screen long enough to read twice.
- Sound effects. A subtle whoosh on a zoom or a pop when text appears adds polish. Keep them quiet.
5. Captions
Generate captions last, on the finished timing (they follow later cuts too, but proofreading once is enough). Choose the style that matches your brand and keep it consistent across videos. For vertical versions, use the Middle position. Captioning in detail →
6. Finish and export
- Colour: a small Temperature and Contrast correction, applied to all clips, makes webcam footage look much better.
- Music: optional for talking heads. If used, keep it very low and instrumental, and tick Lower other sound under voiceovers if you added narration.
- Thumbnail: the AI thumbnail maker picks an expressive frame from the finished edit.
- Export at 1080p, 30 fps, High quality for YouTube.
How long it should take
For a 10-minute raw recording that becomes a 6-minute video, a realistic first edit with this workflow takes 30–60 minutes: a few minutes each for sound, silences and captions, and the rest for the human decisions — structure, hook and B-roll. Those decisions are the part worth your time.
Frequently asked questions
Are jump cuts bad?
No — they're an accepted part of online video. They become tiring only when every sentence jumps. Auto zoom and B-roll disguise many of them.
Should I use a script?
Use bullet points rather than a full script — reading sounds like reading. Remove silences handles the pauses while you think.
How do I make a vertical version?
After exporting the landscape version, set the canvas to 9:16, apply Smart reframe to your clips and export again.