Skip to main content
AI video model
Alibaba
Wan

Wan 2.2 A14B Turbo

Wan 2.2 A14B Turbo is not a single Kyeo workflow. No media selects text-to-video, one image selects image-to-video, and one image plus one audio file selects speech-to-video. Each workflow has different resolutions, credit rates, and visible controls, so confirm the media combination first.

Workflows
Text, image, and image-plus-audio
Credit cost
40/60/80 fixed standard tiers; speech 12/18/24 per second
Prompt limit
Prompt required; up to 5,000 characters
Upload limit
Up to 1 image; speech adds 1 audio file; 10 MB per file
Three model versions are selected dynamically by no media, one image, or one image plus one audio file
Text, image and speech workflows all support 480p, 580p and 720p
Text/image fixed five-second tiers are 40/60/80 credits; speech rates are 12/18/24 credits per output second
The speech workflow sends the current `enable_safety_checker` field with a true default
30-second overview
Wan
Standard 40/60/80; speech 12/18/24 per second
What it does best
Selects text-to-video, image-to-video, or speech-driven video from the supplied media. Each mode uses its own inputs and controls.
Best for
Use it to evaluate prompt-only generation, single-image animation, or image-plus-audio motion as separate tasks. Results from one route should not be generalized to the other two.
Popular searches
How do I use Wan 2.2 A14B Turbo?Which media select each Wan 2.2 workflow?How are the six Wan 2.2 credit prices calculated?

At a glance

What is this model like?

Text and image workflows output a fixed five seconds: 40/60/80 credits at 480p/580p/720p. Speech output seconds equal `num_frames ÷ frames_per_second`, multiplied by 12/18/24 credits for the selected resolution; Kyeo rounds the request total up. Speech requires both image and audio, exposes frame, negative-prompt and inference controls, and sends the current `enable_safety_checker` field with a true default. The page does not promise generation speed, synchronization, image quality, or identity consistency.

What it does best
Selects text-to-video, image-to-video, or speech-driven video from the supplied media. Each mode uses its own inputs and controls.
Best for
Use it to evaluate prompt-only generation, single-image animation, or image-plus-audio motion as separate tasks. Results from one route should not be generalized to the other two.
Why use it on Kyeo AI
Kyeo selects one of the three workflows from the supplied media and shows the standard fixed five-second tiers and speech per-second rates before submission; each route still requires separate result validation.

Key facts

Quickly assess whether this model fits your use case.

Category
AI video model
Vendor
Alibaba
Model family
Wan
Integration
Current site integration
Workflows
Wan — Asynchronous task
Runtime
Asynchronous task
Prompt limit
Kyeo request: non-empty prompt, up to 5,000 characters
Upload limit
JPEG/PNG/WebP; speech adds 1 audio file; 10 MB each

Also known as

The same model may appear under different names across documentation and community discussions; this list keeps them easy to search and compare.

Wan — Which Wan 2.2 workflow exposes a safety check?
Wan — What happens if Wan 2.2 receives audio without an image?
Wan — How should I compare Wan 2.2 with Wan 2.6?

Common questions

These practical questions focus on the task, input conditions, and result requirements you should confirm before choosing.

How do I use Wan 2.2 A14B Turbo?
Which media select each Wan 2.2 workflow?
How are the six Wan 2.2 credit prices calculated?
Do all three Wan 2.2 workflows support 580p?
Which files does Wan 2.2 image-to-video accept?
Does Wan 2.2 speech-to-video require an image and audio?
Why must Wan 2.2 num_frames be divisible by four?
Which Wan 2.2 workflow exposes a safety check?
What happens if Wan 2.2 receives audio without an image?
How should I compare Wan 2.2 with Wan 2.6?

Selection guide

Use these decision points when choosing a model.

1
Choose the workflow from the media first

Use no media for text, one image without audio for image-to-video, or one image plus one audio file for speech.

2
Check resolution and price next

Text and image produce a fixed five seconds for 40/60/80 credits at 480p/580p/720p. Speech costs 12/18/24 credits per output second, calculated from `num_frames ÷ frames_per_second`, and Kyeo rounds the total up.

3
Validate advanced speech controls

For speech mode, set `num_frames` to a multiple of four from 40 to 120 and frame rate from 4 to 60.

Popular comparisons

Compare common alternatives on the same task to clarify differences in inputs, controls, and cost.

How do Wan 2.2 and Wan 2.6 media workflows differ?
Wan 2.2 vs Wan 2.6: resolution and credits
How do Wan 2.2 and Wan 2.5 Video controls differ?
Wan 2.2 vs Wan 2.5 Video: inputs and pricing
How does Wan 2.2 speech mode differ from Infinitalk?
Wan 2.2 vs Infinitalk: image, audio, and prompt contract

Model comparison

Compare the current model with alternatives at a glance.

Media-routed input
Wan 2.2 A14B Turbo
No media, one image, or one image plus one audio file
Wan 2.6
No media, one image, or one video
Infinitalk
Always requires one image, one audio file, and a prompt
Current credits
Wan 2.2 A14B Turbo
Text/image fixed five seconds: 40/60/80; speech: 12/18/24 per output second
Wan 2.6
70-315 by resolution and duration
Wan 2.5 Video
60-200 by resolution and duration
Resolution options
Wan 2.2 A14B Turbo
Text, image and speech all support 480p, 580p and 720p
Wan 2.6
720p or 1080p
Wan 2.5 Video
720p or 1080p
Visible controls
Wan 2.2 A14B Turbo
Text/image use standard controls; speech adds frame, inference and safety controls
Wan 2.6
Resolution and duration only
Infinitalk
Resolution and seed only
Current validation boundary
Wan 2.2 A14B Turbo
Audio needs an image; frame count must divide by four
Wan 2.6
Its video workflow does not use an image or audio file
Infinitalk
Image and audio are always required together

Practical usage insights

Practical guidance based on public sources, current on-site limits, and representative tasks.

current public source workflow names are not Alibaba direct model IDs

Alibaba's official Wan 2.2 A14B release covers text and image models. Alibaba Cloud's current `wan2.2-s2v` also has a separate image-detection step. The current public source A14B speech workflow must be described only from its current public source form and Kyeo payload.

Marketing claims do not guarantee a specific result

720p, 24 fps, fast generation, stable motion, and audiovisual synchronization are not Kyeo result guarantees; use the actual task output.

Capabilities

Text and image modes

Text mode offers prompt, resolution, aspect ratio, prompt expansion, acceleration, and an optional seed. Image mode accepts one reference image instead of a separate aspect-ratio setting and keeps the other visible controls. Safety checking remains enabled.

Speech-driven mode

Speech mode combines a prompt, one image, and one audio clip, with controls for frame count, frame rate, resolution, inference steps, guidance strength, and shift. Negative prompt and seed are optional, and safety checking is enabled by default.

Safety-field boundary

Safety checking is enabled across all modes, but automated checks do not replace manual content-safety and media-rights review.

Use cases

Prompt-only shot checks

Compare 16:9 with 9:16 and 480p/580p/720p, then review subjects, motion, text, and continuity manually.

Single-image animation

Upload one JPEG, PNG, or WebP and describe the desired action and camera motion.

Image-plus-audio motion

Use one supported image and audio file to test speech or body movement. Before publication, review sync, identity, voice, and rights to every input and output.

Prompt tips

Describe visible action first

State the subject, setting, action, and camera movement without phrasing speed, stability, or final quality as guaranteed outcomes.

Do not assume an aspect ratio in image mode

Image mode does not send aspect_ratio. Inspect the source composition, describe the desired motion, and review the result for cropping and edge artifacts.

Separate audio from visual direction

Audio supplies the sound input; the prompt still describes the person's state and visual action. Negative prompt, guidance, and shift are request controls, not synchronization or quality guarantees.

Why choose it

Media presence deterministically selects one of three workflow IDs.
Text and image expose a compact set of aspect, prompt-expansion, seed, and acceleration controls.
Safety checking is enabled by default in every workflow, including speech
All six workflow-specific Kyeo prices and resolutions are visible before submission.

What to know first

Generate one representative sample with the selected settings and confirm that it meets the goal before expanding the batch.
All three workflows support 480p, 580p and 720p, but standard routes use fixed five-second tiers while speech is billed from `num_frames ÷ frames_per_second`.
Speech-mode `num_frames` must be a multiple of four from 40 to 120.
Turbo speed, 24 fps, motion stability, synchronization, image quality, safety filtering, and completion time vary with the input and current public source state and are not Kyeo result guarantees.

Adjustable parameters

Quickly assess whether this model fits your use case.

Resolution
resolution
Optional
Parameter type: select
Default: 720p
480p
580p
720p
Aspect ratio (text only)
aspect_ratio
Optional
Parameter type: select
Default: 16:9
16:9
9:16
Prompt expansion (text/image)
enable_prompt_expansion
Optional
Parameter type: select
Default: false
Off
On
Random seed
seed
Optional
Parameter type: number
Range: 0-2147483647
Acceleration (text/image)
acceleration
Optional
Parameter type: select
Default: none
none
regular
Frame count (speech)
num_frames
Optional
Parameter type: number
Default: 80
Range: 40-120; must be divisible by 4
Frames per second (speech)
frames_per_second
Optional
Parameter type: number
Default: 16
Range: 4-60
Negative prompt (speech)
negative_prompt
Optional
Parameter type: text
Inference steps (speech)
num_inference_steps
Optional
Parameter type: number
Default: 27
Range: 2-40
Guidance scale (speech)
guidance_scale
Optional
Parameter type: number
Default: 3.5
Range: 1-10
Shift (speech)
shift
Optional
Parameter type: number
Default: 5
Range: 1-10
Safety check (speech only)
enable_safety_checker
Optional
Parameter type: select
Default: true
Off
On

Credit usage

Standard 40/60/80; speech 12/18/24 per second

Text and image workflows output a fixed five seconds: 40/60/80 credits at 480p/580p/720p. Speech output seconds equal `num_frames ÷ frames_per_second`, multiplied by 12/18/24 credits for the selected resolution; Kyeo rounds the request total up.

Budget tip

Before batch generation, run an A/B test with the same assets across the current model and alternatives to validate quality and avoid wasting credits.

FAQ

Related models

Compare these similar candidates before deciding.

Wan 2.6

No media selects text-to-video, exactly one image selects image-to-video, and one to three videos select video editing; images and videos cannot be mixed. The default is 1080p, 5 seconds, shot structure Off, and content filtering On.

AI video model
70–315 credits / request

Wan 2.5 Video

Review Wan 2.5 Video on Kyeo AI with text-to-video or one-image input, 5/10 seconds, 720p/1080p, prompt constraints, and 60–200-credit pricing. Kyeo charges by resolution and duration: 60 credits for 720p at 5 seconds, 120 for 720p at 10 seconds, 100 for 1080p at 5 seconds, and 200 for 1080p at 10 seconds. The 720p 5-second default costs 60 credits. How are Wan 2.5's 60/100/120/200 credits calculated? — Review Wan 2.5 Video on Kyeo AI with text-to-video or one-image input, 5/10 seconds, 720p/1080p, prompt constraints, and 60–200-credit pricing.

AI video model
60/100/120/200 credits per request

Infinitalk

Infinitalk uses 1 image, 1 audio file no longer than 15 seconds, and a required prompt to generate a speaking-portrait video on Kyeo. It offers 480p or 720p plus optional `seed`; cost depends on audio duration and resolution rather than a fixed request price.

AI video model
480p: 3 credits/sec; 720p: 12 credits/sec

Sources

Content is based on model documentation, feature references, and the settings available on Kyeo AI.

Source note

Review Wan 2.2 A14B Turbo text, one-image and speech-driven workflows, fixed five-second tiers, per-second speech rates and mode-specific inputs. Text and image workflows output a fixed five seconds: 40/60/80 credits at 480p/580p/720p. Speech output seconds equal `num_frames ÷ frames_per_second`, multiplied by 12/18/24 credits for the selected resolution; Kyeo rounds the request total up.

Last updated: 2026-08-27
Alibaba Cloud: Wan2.2 A14B release
Official
View source
Alibaba Cloud: direct wan2.2-s2v guide
Official
View source