OmniHuman 1.5
The current OmniHuman 1.5 core path turns one image of a person, pet, or animated subject plus one audio clip into a driven video. Signed upload metadata verifies audio duration, which directly determines pre-task authorization.
At a glance
What is this model like?
Images can use any aspect ratio but must still pass true-format and signed-dimension checks. Audio must use a supported format, stay within 10 MB, and be under 60 seconds. The optional prompt follows the stricter 300-character limit rather than a wider conflicting description.
Key facts
Quickly assess whether this model fits your use case.
Also known as
The same model may appear under different names across documentation and community discussions; this list keeps them easy to search and compare.
Common questions
These practical questions focus on the task, input conditions, and result requirements you should confirm before choosing.
Selection guide
Use these decision points when choosing a model.
Use this model when the source is a static person or character image that needs to be driven by audio.
If the original motion and camera need to be preserved, compare Video Lip Sync.
Use a short recommended audio sample to check subject recognition and lip movement before extending the script.
Popular comparisons
Compare common alternatives on the same task to clarify differences in inputs, controls, and cost.
Model comparison
Compare the current model with alternatives at a glance.
Practical usage insights
Practical guidance based on public sources, current on-site limits, and representative tasks.
Limited access is not a complete auxiliary suite
Per-audio-second preauthorization is verifiable
Capabilities
Image-and-audio driving
Two output resolutions
Fast mode
Use cases
Virtual presenter
Animated-character dialogue
Talking-pet clip
Prompt tips
Keep the subject clear
Use clean speech audio
Keep the prompt concise
Why choose it
What to know first
Adjustable parameters
Quickly assess whether this model fits your use case.
Credit usage
Kyeo calculates preauthorization from verified audio duration at 27 credits per second and rounds the total up. Audio must be under 60 seconds, so a valid request is preauthorized for less than 1,620 credits. Final settlement uses actual consumption after completion.
Before batch generation, run an A/B test with the same assets across the current model and alternatives to validate quality and avoid wasting credits.
FAQ
Related models
Compare these similar candidates before deciding.
Kling AI Avatar Standard
Workflow: The UI requires exactly 1 JPEG or PNG image and 1 MPEG, WAV, X-WAV, AAC, MP4, or OGG audio file. The image may be up to 10 MB; audio may be up to 100 MB and 5 minutes. The charge uses the signed audio duration at 8 credits per second and rounds up. A missing, invalid, or over-5-minute duration cannot be submitted.
Infinitalk
Infinitalk uses 1 image, 1 audio file no longer than 15 seconds, and a required prompt to generate a speaking-portrait video on Kyeo. It offers 480p or 720p plus optional `seed`; cost depends on audio duration and resolution rather than a fixed request price.
Volcengine Video Lip Sync
Video Lip Sync rebuilds mouth movement in one existing video to match a target vocal track. Lite mode targets front-facing single-person video, while Basic mode handles more complex single-person scenes and can enable scene detection.
Sources
Content is based on model documentation, feature references, and the settings available on Kyeo AI.
Explore the currently available OmniHuman 1.5 workflow: one image plus one audio clip, 720P or 1080P, audio under 60 seconds, per-second pricing, and limited-access boundaries. Kyeo calculates preauthorization from verified audio duration at 27 credits per second and rounds the total up. Audio must be under 60 seconds, so a valid request is preauthorized for less than 1,620 credits. Final settlement uses actual consumption after completion.
Used to check the input modes, visible controls, media limits, and availability of the version currently connected to this site. The page and workbench show the actual options.