Voiceover-driven auto editing

Turn raw footage into finished short videos — automatically.

Auto Video Editor is a local Python + FFmpeg tool that splits footage, recognizes scenes, assembles clips around your voiceover audio, checks for duplicates and exports with GPU batch encoding — built for TikTok, Douyin and Southeast-Asia e-commerce teams.

100%
local processing
6s
max clip length
GPU
NVENC batch export
9:16
portrait canvas
HOOKClip 1
PAIN POINTClip 2
PRODUCTClip 3
DEMOClip 4
ENDINGClip 5
voiceover.mp3 · hook → pain point → product → demo → ending
Core capabilities

Everything between raw footage and the final cut

A complete pipeline: import → split → label → compose → dedup → preview → export.

📦

Batch material import

Import a whole folder at once. Files are auto-classified into footage and voiceover by name and directory — no manual sorting.

✂️

Shot-change splitting

Footage is split on real shot changes into genuine clips, each saved with a thumbnail and an independent duration.

🧠

Visual AI labeling

Optional CLIP GPU recognition tags scenes and functions — hook, pain point, product, demo, ending and more.

🎙️

Voiceover-driven plans

Each voiceover group generates multiple random plans; the opening prefers the voiceover's own hook frame.

🗂️

Smart dedup checking

Pairwise comparison of clips, fingerprints and per-material usage across all plans, with an automatic retry engine.

🚀

GPU batch export

NVENC/QSV/AMF encoding with transitions, random HSL and speech-rate, verified output and auto cache cleanup.

How it works

A six-stage automated pipeline

Human control stays where it matters — every stage can be reviewed and adjusted.

Import

Select a folder; material is classified into footage and voiceover automatically.

Voiceover

Hooks are recognized and groups are built; the app waits for your confirmation.

Split

Footage is split by shot change into real, thumbnail-backed clips.

Compose

Smart plans map scenes to the voiceover rhythm under strict duration rules.

Review

Drag-and-drop timeline, transitions, preview and dedup reports.

Export

GPU batch export with per-video random parameters and verified output.

Built for

Teams producing short videos at scale

E-commerce

TikTok & Douyin operators

Batch-assemble product showcase, pain-point, demo and ending frames around voiceover audio for dozens of videos at once.

Localization

Multi-language editors

Handles Chinese, Thai, Vietnamese and other content; visuals and audio stay independently adjustable and traceable.

Quality

Teams that need control

Keep manual review and timeline editing rights while automation removes the repetitive, time-consuming work.

Technical specs

Requirements & defaults

Runs fully locally — no cloud upload of your media.

PlatformWindows (local)
RuntimePython 3.12 (64-bit) + FFmpeg
Canvas720×1280 · 9:16
Frame rate30 fps
Encoderh264_nvenc / qsv / amf
Max clip length6 seconds
Default voiceoverup to 90 s / track
Speech-rate variation0.95x – 1.05x
Visual AI (optional)CLIP GPU · ONNX Runtime
Dedup retry threshold35% → 40%
Transitions~70% of suitable boundaries
Export outputMP4 · verified streams
Get started

Up and running on a new computer in minutes

1

Extract

Fully extract the package. Do not run it inside a ZIP archive.

2

Install environment

Double-click 安装环境.cmd — the script installs whatever is missing based on your GPU.

3

Verify

Run 检查环境.cmd to confirm hardware encoding and the optional visual AI pass.

4

Launch

Double-click 启动定制化剪辑工具.cmd.

5

Open

The app opens http://127.0.0.1:8765/ in your browser.

6

Import & go

Drop your material folder in and let the pipeline assemble your videos.

Your footage. Your voiceover. Done for you.

Auto Video Editor runs entirely on your machine — your media never leaves your computer.

Start editing smarter