SmartVidFX Studio | AI Face Replacement, Mandarin Voice Cloning, Lip Sync

Live GPU preview
A livestream host presenting products at a studio desk with a ring light, condenser microphone and phone on a tripod, city skyline behind her
00:14 / 00:38

What this plugin does

SmartVidFX Studio puts a complete AI performance toolkit inside the editor you already use. Drop the effect onto a clip, pick a target identity, and the plugin tracks, replaces and relights the face across the whole shot — no roto, no external app, no round-tripping through a render farm.

The second half of the toolkit is voice. Our speech models are trained on Mandarin first rather than bolted on afterwards, so tones land correctly, numbers and measure words are read the way people actually say them, and a cloned voice keeps its character across takes. Generate a narration track, match it to the new performance, and the lip-sync engine conforms the mouth shapes to the audio automatically.

Built for the way Chinese creators actually work

Short-form teams re-cut the same script for half a dozen platforms every week. Studios localise a finished ad for three regional markets on a two-day turnaround. Both need identity and voice changes that survive close-ups and compression — and both need them without a VFX department. That is the job this plugin was designed around.

  • Local where it can be, cloud where it has to be. Tracking, preview and 1080p replacement run on your machine. The heavier models run on our cluster, which means real upload and download volume — see the section below before you plan a delivery schedule.
  • Frame-accurate tracking. Handles profile turns, partial occlusion, glasses and motion blur without popping.
  • Consent-first identity library. Every stock identity ships with a signed release; your own uploads stay private to your seat.
  • Commercial use included. One licence covers client work, paid ads and platform monetisation for as long as the project lives.

Core modules

Three engines, one timeline

Everything below is included in the same plugin bundle. Enable a module per clip and stack them in any order.

AI face replacement

Swap a performer's identity across a full shot at up to 4K. Skin tone, grain and scene lighting are matched per frame so the result grades like original footage.

Mandarin voice generation

Text-to-speech trained on Chinese, with tone-accurate delivery, adjustable pace and emotion, plus voice cloning from roughly thirty seconds of clean reference audio.

Automatic lip sync

Conform mouth shapes to any generated or recorded track. Re-dub a finished edit into a new language without reshooting a single frame.

Hybrid GPU rendering

Tracking, preview and 1080p work run on your own card and never leave the machine. 4K synthesis and voice cloning run on our cluster, so those jobs transfer your media both ways — budget bandwidth, not just render time.

Batch & versioning

Queue a whole folder of clips against one identity and voice preset, then export platform-specific versions in a single pass.

Provenance & watermarking

Optional invisible watermarking and an exportable edit log, so you can show a client exactly which shots were synthetically modified.

Voice library preview

播音腔 · 男声Broadcast · Male 0:12
亲切口播 · 女声Conversational · Female 0:09
粤语 · 中性Cantonese · Neutral 0:15

Placeholder waveforms — wire these to real audio files before you show this to customers.


Workflow

From raw clip to delivered cut in three moves

No node graphs, no intermediate exports. If you can apply a transition, you can run this.

Drop it on the clip

Drag SmartVidFX from the effects panel onto any clip in the timeline. Tracking analyses the shot in the background while you keep cutting.

Pick an identity and a voice

Choose from the licensed library or upload your own reference. Adjust blend strength, age, and delivery style with a handful of sliders.

Render and ship

Export straight from your NLE. The licence travels with the project, so client deliveries and paid placements are already covered.


Local vs. cloud

Where the processing happens — and how much data moves

Not every workflow runs on your machine. The heavier models run on our inference cluster, which means your source media goes up and the finished render comes back down. Worth understanding before you commit to a delivery date.

Runs on your workstation

  • Face tracking and landmark detection
  • Timeline preview and scrubbing
  • Face replacement up to 1080p
  • Lip-sync conform
  • Relight and grain matching

Nothing leaves the machine for these. Bandwidth is irrelevant.

Runs on our inference cluster

  • 4K face synthesis
  • Voice cloning from your reference audio
  • Dialect and accent voice packs
  • Accelerated batch queue

These upload your source media and download the rendered result.

Plan for the transfer, not just the render

A cloud workflow moves real volume. Ten minutes of 4K ProRes is roughly 30–60 GB going up, and the returned render comes back at a comparable size. On a 100 Mbps uplink that is several hours of transfer before a single frame is processed — frequently longer than the inference itself.

Metered lines, hotel and conference Wi-Fi, VPN tunnels and corporate egress proxies all make this worse, and some will throttle or drop a multi-hour upload outright. If you are working under a data cap, one cloud job can consume it in an afternoon.

Two things help. Proxy mode uploads a low-resolution version, runs inference on that, and applies the resulting transform to your full-quality footage locally — a fraction of the data for most shots. And every cloud job shows an estimated transfer size and duration before it starts, so nothing runs up your connection by surprise.

Everything on the Basic plan runs locally — the cloud-backed workflows listed above are all Pro features, so a Basic licence never transfers media at all.

Cloud-backed steps are labelled in the plugin panel and can be switched off entirely if your footage must not leave the building. You keep every local workflow and lose only the cloud list above.

What's included

  • Face replacement engine with per-frame tracking
  • Mandarin & Cantonese TTS voice pack (24 voices)
  • Voice cloning from a 30-second sample
  • Automatic lip-sync conform
  • Licensed identity library (120 releases)
  • Batch queue and preset manager
  • Relight and skin-grain matching tools
  • Installer for macOS (Apple Silicon) and Windows
  • Bilingual documentation and sample project
  • 12 months of updates on every plan

Works with your editor

  • DR DaVinci Resolve 18 +
  • PR Adobe Premiere Pro
  • AE After Effects
  • FC Final Cut Pro
  • OF OpenFX host
  • CL CLI / watch folder

Names are listed for compatibility only. SmartVidFX Studio is not affiliated with, or endorsed by, any of these vendors.

Tags

More from SmartVidFX Labs

Ready to try it on a real timeline?

Pick a plan, install the plugin, and render your first swap in about ten minutes.