Skip to main content
AI video model
MiniMax
Hailuo

MiniMax H3

MiniMax H3 offers 768P and 2K output with text, first/last-frame, and image/video/audio reference scenes. Reference-video duration and every input image from image 6 onward affect pre-task authorization.

Scene
Text / first-last frames / references
Output
768P / 2K
Duration
4–15 seconds
Credits
64–812
Two output tiers: 768P and 2K
Either a first frame or a last frame can serve as image-mode input
Accepts up to 9 images, 3 videos, and 3 audio clips
Reference-video duration and images from image 6 onward affect preauthorization
30-second overview
Hailuo
64–812 credits
What it does best
A video generation entry point centered on a 2K tier, flexible first/last-frame input, and controlled multimodal references.
Best for
A good fit for high-resolution concept clips, combined character and motion references, first-to-last-frame transitions, and multi-asset short videos.
Popular searches
How do I use MiniMax H3?How much does MiniMax H3 cost?What reference media does MiniMax H3 support?

At a glance

What is this model like?

Image mode requires at least a first frame or a last frame. Reference audio cannot be used alone and must be accompanied by a reference image or video. Every upload has its true format validated, and signed duration, geometry, and FPS are checked before task creation.

What it does best
A video generation entry point centered on a 2K tier, flexible first/last-frame input, and controlled multimodal references.
Best for
A good fit for high-resolution concept clips, combined character and motion references, first-to-last-frame transitions, and multi-asset short videos.
Why use it on Kyeo AI
Kyeo includes media authenticity, count, duration, and additional-image cost in the contract enforced before credits are reserved.

Key facts

Quickly assess whether this model fits your use case.

Reference images
Up to 9
Reference videos
Up to 3; ≤15 seconds total
Reference audio
Up to 3; cannot be used alone
Additional image charge
From image 6 onward: 8 credits each

Also known as

The same model may appear under different names across documentation and community discussions; this list keeps them easy to search and compare.

MiniMax H3 video
Hailuo H3
MiniMax H3 2K video

Common questions

These practical questions focus on the task, input conditions, and result requirements you should confirm before choosing.

How do I use MiniMax H3?
How much does MiniMax H3 cost?
What reference media does MiniMax H3 support?
How do MiniMax H3 768P and 2K differ?
How are MiniMax H3 input images priced?

Selection guide

Use these decision points when choosing a model.

1
For 2K, compare H3

H3 provides a defined 2K tier in this group, with a correspondingly higher per-second rate.

2
Pair audio with a visual reference

A request containing reference audio but no image or video is rejected before task creation.

3
Limit images after the fifth

Each additional image costs 8 credits, so retain only assets that contribute distinct information.

Popular comparisons

Compare common alternatives on the same task to clarify differences in inputs, controls, and cost.

How do the resolutions and rates of MiniMax H3 and Seedance 2.0 Mini differ?
How do the reference-media capabilities of MiniMax H3 and HappyHorse 1.1 compare?

Model comparison

Compare the current model with alternatives at a glance.

Reference mode
MiniMax H3
9 images / 3 videos / 3 audio clips; 2K available
Seedance 2.0 Mini
9 images / 3 videos / 3 audio clips; up to 720p
HappyHorse 1.1
Up to 9 images; no reference video or audio

Practical usage insights

Practical guidance based on public sources, current on-site limits, and representative tasks.

Media count directly affects the cost formula

H3 is not priced solely by seconds; every input image from image 6 onward also increases preauthorization.

The narrower format range is a product-safety boundary

Kyeo enables only JPEG, PNG, and WebP formats that the current path can identify reliably, rather than presenting unverified formats as usable.

Capabilities

2K video generation

Choose 2K for text, image, or reference scenes.

Multimodal references

Combine up to 9 images, 3 videos, and 3 audio clips.

First- or last-frame input

Use a first frame, a last frame, or both in image mode.

Use cases

High-resolution concept clip

Use the 2K tier to test a brand-film or visual-proposal direction.

Combined motion and character reference

Assign separate roles to a character image and a motion video.

Build toward a target last frame

Upload only the desired final frame and describe how the shot should arrive there.

Prompt tips

Refer to visual elements explicitly

Explain how the person, clothing, or setting in each image should appear in the video.

Align audio with action

State when the sound should occur and which action it accompanies.

Control detail at 2K

Prioritize subject, material, lighting, and camera direction instead of stacking conflicting style terms.

Why choose it

Provides a defined 2K output tier.
Flexible first- and last-frame input.
Clear multimodal-reference limits.
Additional-image charges are visible before task creation.

What to know first

Output is limited to 15 seconds.
Reference audio cannot be used alone.
Images from image 6 onward increase the cost.
Kyeo does not enable image formats that the current path cannot identify reliably.

Adjustable parameters

Quickly assess whether this model fits your use case.

Generation scene
seedance_scene
Optional
Parameter type: Select
Default: text
Text
First frame / last frame
Multimodal references
Resolution
resolution
Optional
Parameter type: Select
Default: 2K
768P
2K
Video duration
duration
Required
Parameter type: Integer
Default: 6 seconds
4–15 seconds

Credit usage

64–812 credits

768P and 2K cost 16 and 26 credits per second, multiplied by output duration plus verified total reference-video duration. The first 5 input images add no separate charge; images 6–9 add 8 credits each. Kyeo rounds the total up before task creation.

Budget tip

Before batch generation, run an A/B test with the same assets across the current model and alternatives to validate quality and avoid wasting credits.

FAQ

Related models

Compare these similar candidates before deciding.

Seedance 2.0 Mini

Seedance 2.0 Mini combines text, first/last-frame, and multimodal-reference generation in one entry point. Reference video enters the pre-task pricing formula, so both media duration and output duration must be verified before task creation.

AI video model
16–150 credits

Seedance 2.5

Seedance 2.5 combines a 30-second boundary, higher media counts, and video-editing reference input in one strict scene contract. Automatic duration is not underestimated at the 5-second default; Kyeo preauthorizes against the 30-second maximum.

AI video model
102–4,110 credits

HappyHorse 1.1

HappyHorse 1.1 routes by image count: no image selects text mode, 1 image selects first-frame image-to-video, and 2–9 images select reference mode. All three modes share the same resolutions, durations, and per-second rates.

AI video model
68–435 credits

Sources

Content is based on model documentation, feature references, and the settings available on Kyeo AI.

Source note

Explore MiniMax H3: 768P or 2K, 4–15 seconds, first/last frames, 9 images, 3 videos, and 3 audio clips, 64–812-credit preauthorization, and media limits. 768P and 2K cost 16 and 26 credits per second, multiplied by output duration plus verified total reference-video duration. The first 5 input images add no separate charge; images 6–9 add 8 credits each. Kyeo rounds the total up before task creation.

Last updated: 2026-08-28
MiniMax H3 video current interface notes
Platform

Used to check the input modes, visible controls, media limits, and availability of the version currently connected to this site. The page and workbench show the actual options.