Skip to main content
AI video model
Volcengine
Video Lip Sync

Volcengine Video Lip Sync

Video Lip Sync rebuilds mouth movement in one existing video to match a target vocal track. Lite mode targets front-facing single-person video, while Basic mode handles more complex single-person scenes and can enable scene detection.

Workflow
Video + audio lip sync
Modes
Lite / Basic
Video
360p–1080p; 24–60 FPS
Pricing
8 / audio second
A clear existing-video-plus-target-audio workflow
Lite and Basic modes for two kinds of single-person footage
Upfront validation of video resolution, FPS, bitrate, and file size
Preauthorization based on signed audio seconds
30-second overview
Video Lip Sync
8 credits per audio second
What it does best
A target-audio lip-redirection workflow for existing video, rather than a character-video generator built from a still image.
Best for
A good fit for replacing dubbing, multilingual talking-head clips, single-person explainers, and content that must retain the original movement and camera work.
Popular searches
How do I use Video Lip Sync?How much does AI video lip sync cost?How do I sync lips after dubbing a video?

At a glance

What is this model like?

Video must meet 360p–1080p, 24–60 FPS, and 1–30 Mbps boundaries, plus Kyeo's 95 MB file limit. Signed upload metadata proves file size, duration, resolution, and FPS, while verified audio duration directly determines pre-task authorization.

What it does best
A target-audio lip-redirection workflow for existing video, rather than a character-video generator built from a still image.
Best for
A good fit for replacing dubbing, multilingual talking-head clips, single-person explainers, and content that must retain the original movement and camera work.
Why use it on Kyeo AI
Kyeo verifies true video metadata and audio duration before credits are reserved and strictly separates parameters that are exclusive to each mode.

Key facts

Quickly assess whether this model fits your use case.

Video formats
MP4 / MOV
Video bitrate
1–30 Mbps
Audio
Clean target vocal; 10 MB
Output duration
Follows the audio
Output
MP4 at 25 FPS; duration follows the target audio

Also known as

The same model may appear under different names across documentation and community discussions; this list keeps them easy to search and compare.

Volcengine video lip sync
AI video lip sync
video dubbing lip synchronization

Common questions

These practical questions focus on the task, input conditions, and result requirements you should confirm before choosing.

How do I use Video Lip Sync?
How much does AI video lip sync cost?
How do I sync lips after dubbing a video?
How do Video Lip Sync Lite and Basic modes differ?
What videos does Video Lip Sync support?

Selection guide

Use these decision points when choosing a model.

1
Start with Lite mode for a front-facing subject

Lite mode fields fit straightforward footage with one front-facing person and simple looping alignment.

2
Use Basic mode for a complex single-person scene

Choose Basic mode when the workflow needs scene segmentation and speaker-recognition direction.

3
Switch to a driven-video model for a still image

This model requires video input. When only a still image is available, compare OmniHuman 1.5.

Popular comparisons

Compare common alternatives on the same task to clarify differences in inputs, controls, and cost.

How do the source-media requirements of Video Lip Sync and OmniHuman 1.5 differ?
When should I choose Lite mode or Basic mode?

Model comparison

Compare the current model with alternatives at a glance.

Source visual
Volcengine Video Lip Sync
Existing video + target audio
OmniHuman 1.5
Still image + target audio
Kling AI Avatar Standard
Still image + target audio

Practical usage insights

Practical guidance based on public sources, current on-site limits, and representative tasks.

Lip sync is not a complete dubbing workflow

You still need a clean target track that you have permission to use, followed by human review of meaning, pacing, and visual synchronization.

Trusted media metadata supports both safety and pricing

Signed video metadata enforces the media contract, while signed audio duration drives pricing; neither can be self-reported by the client.

Capabilities

Video lip redirection

Preserve the source video's visuals while target audio drives mouth movement.

Lite looping alignment

Loop the video when audio is longer and optionally use reverse looping.

Complex-scene handling

Enable scene detection and speaker recognition in Basic mode.

Use cases

Multilingual talking-head video

Replace an explainer video's audio with another language and synchronize the mouth movement.

Course narration update

Keep the instructor's motion and visuals while replacing the narration track.

Marketing-video revision

Replace a short spoken segment without reshooting the main footage.

Prompt tips

Prepare clean vocal audio

Reduce music, reverberation, and overlapping speakers so the driving signal stays clear.

Choose a clear front-facing shot

Lite mode works best with one front-facing person, a visible mouth, and little occlusion.

Check audio and video lengths first

Choose looping, reverse looping, or a template start based on the mode to avoid unnatural repetition.

Why choose it

Retains the camera work and motion of existing video.
Low per-second rate with precise pre-task authorization.
Two modes cover simple and complex single-person scenes.
Signed evidence validates critical video metadata.

What to know first

Designed only for single-person scenes.
Requires both a video and target audio.
Kyeo limits video files to 95 MB.
Automated lip sync still needs human review for misalignment, expression, and semantic naturalness.

Adjustable parameters

Quickly assess whether this model fits your use case.

Service mode
mode
Required
Parameter type: Select
Default: lite
Lite
Base
Vocal separation
separate_vocal
Optional
Parameter type: Toggle
Default: off
Off
On
Scene detection and speaker recognition
open_scenedet
Base mode only
Parameter type: Toggle
Default: off
Off
On
Loop video when audio is longer
align_audio
Lite mode only
Parameter type: Toggle
Default: on
On
Off
Reverse loop
align_audio_reverse
Lite mode only
Parameter type: Toggle
Default: off
Off
On
Template start time
templ_start_seconds
Lite mode only
Parameter type: Number
Default: 0 seconds

Credit usage

8 credits per audio second

Kyeo calculates preauthorization at 8 credits per verified second of target audio and rounds the total up. Output duration follows the audio, and final settlement uses actual consumption after completion.

Budget tip

Before batch generation, run an A/B test with the same assets across the current model and alternatives to validate quality and avoid wasting credits.

FAQ

Related models

Compare these similar candidates before deciding.

OmniHuman 1.5

The current OmniHuman 1.5 core path turns one image of a person, pet, or animated subject plus one audio clip into a driven video. Signed upload metadata verifies audio duration, which directly determines pre-task authorization.

AI video model
27 credits per audio second

Kling AI Avatar Standard

Workflow: The UI requires exactly 1 JPEG or PNG image and 1 MPEG, WAV, X-WAV, AAC, MP4, or OGG audio file. The image may be up to 10 MB; audio may be up to 100 MB and 5 minutes. The charge uses the signed audio duration at 8 credits per second and rounds up. A missing, invalid, or over-5-minute duration cannot be submitted.

AI video model
8 credits / audio second

Infinitalk

Infinitalk uses 1 image, 1 audio file no longer than 15 seconds, and a required prompt to generate a speaking-portrait video on Kyeo. It offers 480p or 720p plus optional `seed`; cost depends on audio duration and resolution rather than a fixed request price.

AI video model
480p: 3 credits/sec; 720p: 12 credits/sec

Sources

Content is based on model documentation, feature references, and the settings available on Kyeo AI.

Source note

Explore Video Lip Sync: video plus audio input, Lite and Basic modes, 360p–1080p, 24–60 FPS, 8 credits per audio second, and parameter boundaries. Kyeo calculates preauthorization at 8 credits per verified second of target audio and rounds the total up. Output duration follows the audio, and final settlement uses actual consumption after completion.

Last updated: 2026-08-28
Volcengine video lip sync current interface notes
Platform

Used to check the input modes, visible controls, media limits, and availability of the version currently connected to this site. The page and workbench show the actual options.