MiniMax H3
MiniMax H3 offers 768P and 2K output with text, first/last-frame, and image/video/audio reference scenes. Reference-video duration and every input image from image 6 onward affect pre-task authorization.
At a glance
What is this model like?
Image mode requires at least a first frame or a last frame. Reference audio cannot be used alone and must be accompanied by a reference image or video. Every upload has its true format validated, and signed duration, geometry, and FPS are checked before task creation.
Key facts
Quickly assess whether this model fits your use case.
Also known as
The same model may appear under different names across documentation and community discussions; this list keeps them easy to search and compare.
Common questions
These practical questions focus on the task, input conditions, and result requirements you should confirm before choosing.
Selection guide
Use these decision points when choosing a model.
H3 provides a defined 2K tier in this group, with a correspondingly higher per-second rate.
A request containing reference audio but no image or video is rejected before task creation.
Each additional image costs 8 credits, so retain only assets that contribute distinct information.
Popular comparisons
Compare common alternatives on the same task to clarify differences in inputs, controls, and cost.
Model comparison
Compare the current model with alternatives at a glance.
Practical usage insights
Practical guidance based on public sources, current on-site limits, and representative tasks.
Media count directly affects the cost formula
The narrower format range is a product-safety boundary
Capabilities
2K video generation
Multimodal references
First- or last-frame input
Use cases
High-resolution concept clip
Combined motion and character reference
Build toward a target last frame
Prompt tips
Refer to visual elements explicitly
Align audio with action
Control detail at 2K
Why choose it
What to know first
Adjustable parameters
Quickly assess whether this model fits your use case.
Credit usage
768P and 2K cost 16 and 26 credits per second, multiplied by output duration plus verified total reference-video duration. The first 5 input images add no separate charge; images 6–9 add 8 credits each. Kyeo rounds the total up before task creation.
Before batch generation, run an A/B test with the same assets across the current model and alternatives to validate quality and avoid wasting credits.
FAQ
Related models
Compare these similar candidates before deciding.
Seedance 2.0 Mini
Seedance 2.0 Mini combines text, first/last-frame, and multimodal-reference generation in one entry point. Reference video enters the pre-task pricing formula, so both media duration and output duration must be verified before task creation.
Seedance 2.5
Seedance 2.5 combines a 30-second boundary, higher media counts, and video-editing reference input in one strict scene contract. Automatic duration is not underestimated at the 5-second default; Kyeo preauthorizes against the 30-second maximum.
HappyHorse 1.1
HappyHorse 1.1 routes by image count: no image selects text mode, 1 image selects first-frame image-to-video, and 2–9 images select reference mode. All three modes share the same resolutions, durations, and per-second rates.
Sources
Content is based on model documentation, feature references, and the settings available on Kyeo AI.
Explore MiniMax H3: 768P or 2K, 4–15 seconds, first/last frames, 9 images, 3 videos, and 3 audio clips, 64–812-credit preauthorization, and media limits. 768P and 2K cost 16 and 26 credits per second, multiplied by output duration plus verified total reference-video duration. The first 5 input images add no separate charge; images 6–9 add 8 credits each. Kyeo rounds the total up before task creation.
Used to check the input modes, visible controls, media limits, and availability of the version currently connected to this site. The page and workbench show the actual options.