
Wan 2.6 Reference-to-Video Spicy
Wan 2.6 Reference-to-Video Spicy is Alibaba's advanced uncensored model designed for high-fidelity, reference-guided video synthesis. Operating as a Spicy uncensored model, it completely bypasses standard safety guardrails to perform dynamic character replication, motion transfer, and style transformation without content restrictions. By accurately tracking facial features, body postures, and spatial trajectories from reference images or video clips, it generates fluid, high-amplitude action sequences while preserving key visual identity elements. Wan 2.6 Reference-to-Video Spicy efficiently powers unrestricted choreography, motion transfer pre-visualization, digital human performances, and high-impact cinematic visual effects.
Read Me
Wan 2.6 Reference-to-Video Spicy API
Wan 2.6 Reference-to-Video Spicy is Alibaba Cloud's next-generation reference-guided video generation model, released with the Wan 2.6 series on December 16, 2025. Per Alibaba Cloud's official announcement, it is China's first reference-to-video generation model: upload character reference footage containing appearance and voice, then use a text prompt to place that same character in entirely new scenes. Each run outputs up to 15 seconds at 720P or 1080P, with consistent identity preservation and precise action alignment.
The model is offered through the iCreat platform as an API using a two-step asynchronous task flow: submit a task to obtain a task_id, then poll the result endpoint until completion. It supports 720P and 1080P resolutions (four landscape and portrait sizes), billed by output video duration: $0.10 per second at 720P and $0.15 per second at 1080P, with reference media input not billed separately.
Model Positioning
Wan 2.6 Reference-to-Video Spicy targets video production where character consistency comes first. Unlike text-to-video generation from scratch, it anchors the subject's appearance and timbre with reference media, then drives the performance with a prompt — suited to continuous content that features the same character across scenes. Within the Wan 2.6 family it complements text-to-video and image-to-video: text-to-video starts free creation from text, image-to-video animates stills, and reference-to-video casts existing characters into new scenes. Typical needs include character consistency content, style transfer, e-commerce showcases, and creative visual production.
Core Capabilities
Reference-Guided Generation
Generates dynamic, high-fidelity video conditioned on external reference media such as reference video clips, character identity references, and scene elements. The prompt describes the target scene, character actions, and details, while reference media anchors subject appearance and style.
Identity and Timbre Preservation
Transfers the reference subject's appearance and timbre into newly generated scenes. People, animals, and objects can all take the lead role, and multiple subjects can co-star while keeping their distinct identities, with realistic motion dynamics.
HD Output with Flexible Duration
Supports 720P and 1080P resolutions, mapping to four landscape and portrait sizes: 1280*720, 720*1280, 1920*1080, and 1080*1920. Each run generates 4–15 seconds, covering the mainstream short-video duration range.
Precise Action Alignment
Aligns actions and framing precisely with the prompt, so character performances and camera work match the described narrative intent.
Pricing
| Resolution | Unit Price (USD/Second) | Cost for 5 Seconds |
|---|---|---|
| 720P | 0.1 | 0.5 |
| 1080P | 0.15 | 0.75 |
Total Cost = Unit Price × Output Video Duration. Only the output video duration is billed; reference media input is not charged separately. The costUSD field in the response returns the actual cost once the task succeeds.
Application Scenarios
- Character-consistency content: IP characters, creators, and virtual personas appearing stably across scenes
- Short dramas and narrative creation: role play, multi-subject co-starring, and continuous storyline production
- E-commerce showcases: display and seeding videos generated from product and model reference media
- Style transfer: migrating the look and style of reference media into entirely new scenes
- Creative visual production: rapid production of advertising, promotional, and social media videos
Model Comparison
Same-Series Comparison
| Model | Task Type | Input Media | Max Duration | Resolution | Audio Capability |
|---|---|---|---|---|---|
| Wan 2.6 Reference-to-Video Spicy (this model) | Reference-to-video | Text + reference video/images | 15 seconds | 720P / 1080P | Preserves character timbre |
| Wan 2.6 Text-to-Video | Text-to-video | Text | 15 seconds | 720P / 1080P | Native joint audio-video generation |
| Wan 2.6 Image-to-Video | Image-to-video | Text + first-frame image | 15 seconds | 720P / 1080P | Native joint audio-video generation |
Cross-Model Comparison
| Model | Developer | Task Type | Input Media | Duration | Audio Capability |
|---|---|---|---|---|---|
| Wan 2.6 Reference-to-Video Spicy | Alibaba Cloud | Reference-to-video | Text + reference video/images | 4–15 seconds | Preserves character timbre |
| MiniMax H3 | MiniMax | Text-to-video / reference-to-video | Text / images / video / audio | 5–15 seconds | Native stereo audio |
| Seedance 2.0 | ByteDance | Text-to-video | Text | 4–15 seconds | Audio generation |
| Seedance 2.0 Mini | ByteDance | Text-to-video | Text | 4–15 seconds | Audio generation |
Same-series data comes from Tongyi Wanxiang's official release; this model's specifications follow the iCreat channel documentation.
Why Choose Wan 2.6 Reference-to-Video Spicy?
- Reference media anchors character identity, so prompts no longer need to re-describe every character detail — less effort for continuous content production
- Dual consistency of appearance and timbre: the character not only looks right on screen, the voice travels with the character
- Transparent billing: charged by output video duration, with reference media input not billed
- Two-step async interface with clear status semantics;
statusandcostUSDmake task progress and cost transparent - 720P / 1080P HD output with 4–15 second durations covering mainstream short-video scenarios
Specifications
| Item | Description |
|---|---|
| Base URL | https://api.icreat.ai |
| Submit endpoint | POST /v1/task/submit/aliyun/wan2-6/reference-to-video-global |
| Query result endpoint | POST /v1/task/result |
| Model ID | aliyun/wan2-6/reference-to-video-global |
| Authentication | Authorization: Bearer header |
| Call pattern | Two-step asynchronous task (submit → poll) |
input.prompt |
Required, string, describing the video to generate, character actions, and scene details |
input.reference_urls |
Required, array of HTTPS URLs pointing to reference videos or images |
parameters.size |
Required, 1280*720 / 720*1280 / 1920*1080 / 1080*1920 |
parameters.duration |
Required, integer, 4–15 |
| Task status | SUBMITTED / SUCCEEDED / FAILED |
| Output resource | type is Video, includes url and download_url |
| Cost field | costUSD, present only on SUCCEEDED |
Architecture
The Wan 2.6 series was released in December 2025, covering five model types: text-to-video, image-to-video, reference-to-video, image generation, and text-to-image. Reference-to-video uses a reference-guided generation architecture: the model conditions on external reference media (reference video clips, character identity references, scene elements), transfers the reference subject's appearance and timbre into new scenes, and generates dynamic footage from the text prompt while preserving identity consistency across multiple co-starring subjects. At the series level, Wan 2.6 supports multi-shot narrative and native audio-video sync, with improved instruction following and visual quality over the previous generation.
Notes
- Submit and query must be chained with the same
task_id - While processing,
resultis[]; keep polling, with a suggested interval of 2–5 seconds FAILEDis terminal; check request parameters and reference media URLs before resubmitting- URLs in
reference_urlsmust be HTTPS inputandparametersare sibling top-level fields; do not nest thempromptandreference_urlsare both required; omitting either blocks submissioncostUSDis returned only when the task succeeds; failed tasks do not produce a cost field
FAQ
How do I get started with the Wan 2.6 Reference-to-Video Spicy API?
Register on the iCreat platform and obtain an API Key from the console, send a generation request to the submit endpoint, then poll the result endpoint with the returned task_id. All requests are authenticated with the Authorization: Bearer header. Keep your API Key safe and never expose it in client code or public repositories.
How is the cost calculated?
Total Cost = Unit Price × Output Video Duration. 720P is $0.10 per second and 1080P is $0.15 per second; only the output duration is billed and reference media input is not charged separately. For example, a 720P, 5-second video costs $0.50. The costUSD field in the response gives the actual cost once the task succeeds.
What reference media is supported?
Reference videos and images are supported, passed as an array of HTTPS URLs via reference_urls to guide character appearance, motion, or style. Prompts can use names like Character1 and Character2 to refer to different reference subjects and describe their actions and interactions.
What resolutions and aspect ratios are supported?
parameters.size accepts 1280*720, 720*1280, 1920*1080, and 1080*1920, corresponding to landscape and portrait formats of 720P and 1080P resolutions, so you can pick the right size for each delivery platform.
How do I query the task result?
Send a POST request to https://api.icreat.ai/v1/task/result with the task_id in the body. A status of SUBMITTED means processing, with result as [] — keep polling. When status is SUCCEEDED, result returns the video resource array; read url or download_url to retrieve the video.
What if the task fails?
FAILED is a terminal status. Check in order: whether the prompt is empty, whether reference media URLs are publicly accessible, whether size is one of the four allowed values, and whether duration is within 4–15. Fix any issue and resubmit. The cost field is returned only when the task succeeds.

