Volcengine Video Lip Sync
Video Lip Sync rebuilds mouth movement in one existing video to match a target vocal track. Lite mode targets front-facing single-person video, while Basic mode handles more complex single-person scenes and can enable scene detection.
At a glance
What is this model like?
Video must meet 360p–1080p, 24–60 FPS, and 1–30 Mbps boundaries, plus Kyeo's 95 MB file limit. Signed upload metadata proves file size, duration, resolution, and FPS, while verified audio duration directly determines pre-task authorization.
Key facts
Quickly assess whether this model fits your use case.
Also known as
The same model may appear under different names across documentation and community discussions; this list keeps them easy to search and compare.
Common questions
These practical questions focus on the task, input conditions, and result requirements you should confirm before choosing.
Selection guide
Use these decision points when choosing a model.
Lite mode fields fit straightforward footage with one front-facing person and simple looping alignment.
Choose Basic mode when the workflow needs scene segmentation and speaker-recognition direction.
This model requires video input. When only a still image is available, compare OmniHuman 1.5.
Popular comparisons
Compare common alternatives on the same task to clarify differences in inputs, controls, and cost.
Model comparison
Compare the current model with alternatives at a glance.
Practical usage insights
Practical guidance based on public sources, current on-site limits, and representative tasks.
Lip sync is not a complete dubbing workflow
Trusted media metadata supports both safety and pricing
Capabilities
Video lip redirection
Lite looping alignment
Complex-scene handling
Use cases
Multilingual talking-head video
Course narration update
Marketing-video revision
Prompt tips
Prepare clean vocal audio
Choose a clear front-facing shot
Check audio and video lengths first
Why choose it
What to know first
Adjustable parameters
Quickly assess whether this model fits your use case.
Credit usage
Kyeo calculates preauthorization at 8 credits per verified second of target audio and rounds the total up. Output duration follows the audio, and final settlement uses actual consumption after completion.
Before batch generation, run an A/B test with the same assets across the current model and alternatives to validate quality and avoid wasting credits.
FAQ
Related models
Compare these similar candidates before deciding.
OmniHuman 1.5
The current OmniHuman 1.5 core path turns one image of a person, pet, or animated subject plus one audio clip into a driven video. Signed upload metadata verifies audio duration, which directly determines pre-task authorization.
Kling AI Avatar Standard
Workflow: The UI requires exactly 1 JPEG or PNG image and 1 MPEG, WAV, X-WAV, AAC, MP4, or OGG audio file. The image may be up to 10 MB; audio may be up to 100 MB and 5 minutes. The charge uses the signed audio duration at 8 credits per second and rounds up. A missing, invalid, or over-5-minute duration cannot be submitted.
Infinitalk
Infinitalk uses 1 image, 1 audio file no longer than 15 seconds, and a required prompt to generate a speaking-portrait video on Kyeo. It offers 480p or 720p plus optional `seed`; cost depends on audio duration and resolution rather than a fixed request price.
Sources
Content is based on model documentation, feature references, and the settings available on Kyeo AI.
Explore Video Lip Sync: video plus audio input, Lite and Basic modes, 360p–1080p, 24–60 FPS, 8 credits per audio second, and parameter boundaries. Kyeo calculates preauthorization at 8 credits per verified second of target audio and rounds the total up. Output duration follows the audio, and final settlement uses actual consumption after completion.
Used to check the input modes, visible controls, media limits, and availability of the version currently connected to this site. The page and workbench show the actual options.