Wan AI Video Generator: Try Wan 3.0 Online, with Sound
Make Wan AI videos in your browser with Wan 3.0, the newest Wan model from Alibaba. Type a prompt or add an image and get one continuous take of up to 30 seconds at up to 1080p, with dialogue, effects and ambience made in the same pass. No Alibaba Cloud account, no GPU, nothing to install.
- 100 free credits
- Up to 30s in one take
- Native 1080p with sound
- Text or image to video
Sexually explicit (NSFW) content is strictly prohibited.
Wan 3.0 examples
Clips made with Wan 3.0, sound included, each with the prompt behind it. Hover to play, turn the sound on, open the prompt — or put it in the box and make your own version.
Long single takes, up to 30 seconds
Each one is a single generation, not shots cut together.
30 seconds in one take
An eagle’s-eye flight from a snowy cliff past a saber-toothed cat and fighting giant elk to an erupting volcano, without a cut.
Show the prompt
Global Style & Setting: Photorealistic epic nature documentary style (like Prehistoric Planet or Planet Earth). 30 seconds of CG epic one-shot. We follow the FPV perspective of a Haast's Eagle (the largest eagle in history, with a 3-meter wingspan), flying continuously from a snowy cliff nest, across the Pleistocene glacial forests, interacting thrillingly with prehistoric megafauna. The camera blends high-speed flight, elegant gliding, thrilling dodges, and grand reveals. Lighting transitions from the warm golden hour of sunset to the cold blue of glacial twilight, ending with the fiery red of a volcanic eruption. [00.00-07.00s] Departure & Gliding: The Snowy Cliff Prompt: [FPV POV, One continuous shot]. The camera starts in a massive nest made of rocks and dry branches on a snowy cliff edge. Several fluffy Haast's eagle chicks are screeching beside us. With a high-pitched cry, we (the eagle) spread our massive, feathered wings, take a few running steps, and leap off the cliff. The camera dives into an extreme, heart-pounding freefall. After a moment of weightlessness, we spread our wings, catching the updraft to glide smoothly. The camera transitions to a steady, forward-tracking climb. Below us is the turbulent, icy blue ocean, with the setting sun casting shimmering golden light on the waves. Details: We can see the primary feathers trembling in the wind. The sound of howling wind mixes with the crashing of waves against the ice. [07.00-14.00s] Skimming the Ice & The Ambush Prompt: We drop altitude, flying almost skimming the surface of a frozen glacial lake. The camera arcs right, counter-clockwise, tracking the shoreline. On the ice, a herd of giant, flightless Moas (terror birds) are foraging. Suddenly, a massive shadow bursts from the snowdrifts—a Smilodon (saber-toothed cat) leaps out, ambushing a lone Moa. A massive explosion of snow and ice shards erupts, almost splashing the lens. We react instantly, pulling up sharply in an emergency ascent, narrowly dodging the flying ice. Lighting/Performance: The Moas scatter in panic. The Smilodon sinks back into the snow. The setting sun refracts through the suspended ice mist, creating a brief, stunning ice-rainbow. [14.00-23.00s] Forest Chase & The Antler Gap Escape Prompt: Flying over the ice, we dive into a dense, lush forest of prehistoric Araucaria trees. To dodge a rival, more aggressive Haast's Eagle patrolling the canopy, we plunge beneath the tree line. The camera switches to extreme high-speed FPV drone mode, weaving rapidly between giant ferns and towering trunks. Suddenly, we burst from behind a tree into a clearing. Two colossal Megaloceros (Irish Elks with 12-foot antlers) are locking horns in a violent, earth-shaking battle. [The Antler Gap穿越] We try to fly around, but the elks charge directly into our path! Time seems to slow down. The camera executes a dolly zoom (Vertigo effect); the forest background rushes backward, while the massive, interlocking, tree-branch-like antlers seem to fill the entire world, crushing toward us! With nowhere to go, we choose the only path—flying directly through the narrow, shifting gap between their crashing antlers! Details: The camera猛地 (suddenly) dives into the gap. For 0.5 seconds, the frame is surrounded by the rough, moss-covered texture of the giant antlers, snapping twigs, and flying snow. Just milliseconds before the antlers crush together with a deafening "CRACK," we slip out through the bottom gap! The camera shakes violently from the shockwave, and a few snowflakes stick to the lens. [23.00-30.00s] Soaring Over Glaciers & The Epic Panorama Prompt: Escaping by a hair's breadth, we pull up frantically, gaining altitude to leave the danger below. The Megaloceros let out a furious, bellowing roar. Riding the thermal currents, we begin a spiral climb. As our altitude increases, the entire glacial archipelago is revealed. The camera pulls up to reveal a distant volcano erupting, with lava flowing like red rivers down the black rock. Below on the frozen plains, thousands of Moas are migrating in a massive herd. Finally, exhausted but reborn, we fly into the dark red clouds stained by volcanic ash and the glowing Aurora Borealis, leaving only a silhouette. The screen fades to black. Lighting/Emotion: The entire world is dyed in a magnificent, apocalyptic red by the sunset and the volcano. The sound of our heavy, post-survival breathing mixes with the wind, adding a sense of life's fragility and resilience to the grand finale.
Desert chase, 20 seconds
A drone follows an SUV through a red-rock canyon as a sandstorm swallows the ruins behind it.
Show the prompt
[One-take, FPV drone chase tracking shot, no cuts, 20-second long take, post-apocalyptic wasteland cinematic feel] Overall parameters: 20 seconds, 4K 60fps, 16:9 landscape. Realistic post-apocalyptic wasteland style; epic lighting, physics-level dust particle effects, realistic vehicle dynamics and suspension details, warm golden wasteland tones, shallow depth of field, extreme car body paint and material details. Global setting: A post-apocalyptic red rock canyon, with weathered ancient cave ruins on the cliff walls on both sides, low-angle golden afternoon sunlight. On the distant horizon, a dark-golden sandstorm wall several thousand meters high is advancing, swallowing everything in its path. The subject is a silver-white futuristic luxury SUV, streamlined body, closed grille, no brand logos or text anywhere on the vehicle, body covered with a thin layer of dust, driverless, moving autonomously. [0–3 seconds: Wheel close-up start] Camera movement: Extremely low-angle close-up, camera close to the ground aimed at the wheels. Frame: The SUV’s wheels roll over the sandy ground, slowly starting and accelerating, dust kicked up from the tires, with the mechanical details of the suspension and wheel hubs clearly visible. In the distant background behind the wheels, the sandstorm wall presses down on the horizon like a sky curtain, advancing toward the canyon. [3–7 seconds: Low-altitude side tracking shot · chase] Camera movement: Low-altitude parallel side tracking, camera moving sideways at the same speed as the vehicle. Frame: The SUV drives at high speed through the canyon, paint surface shimmering with reflected light, the rear dragging a long golden dust trail. The sandstorm wall is visibly closer—it swallows the cliff walls and cave ruins it passes, ancient caves disappearing one by one in the sandstorm, the unswallowed cliff walls on both sides receding rapidly, maximizing the oppressive feeling of the approaching storm. [7–10 seconds: Through-window shot · eerie] Camera movement: Fly-through shot, camera accelerates to catch up with the vehicle, smoothly passes through the rear window into the cockpit, sweeps across the interior, then exits through the windshield. Frame: The camera passes through the cabin—an empty driverless cockpit, wraparound ambient lighting and the horizontal large screen faintly glowing, leather and metal texture details visible; the instant it exits through the windshield, the sandstorm wall has already filled the entire far end of the forward view, with a faint rumbling sound coming from afar. [10–14 seconds: Drift orbit · giant wreckage] Camera movement: Orbit shot, the SUV drifts with its rear swinging out, camera orbits 180 degrees around the vehicle as the center. Frame: Ahead on the sandy ground, a hundreds-of-meters-tall giant mechanical wreckage is diagonally embedded (a half-weathered mechanical arm, covered in rust and dust). The SUV drifts around the wreckage, body sliding sideways, dust violently rising in a ring and covering half the frame, sunlight piercing through the sand curtain forming a golden halo, the shadow of the wreckage and the sandstorm wall overlapping in the background. [14–17 seconds: Cliff leap · slow motion] Camera movement: Vertical rise followed by slow motion, camera lifts with the vehicle as it becomes airborne. Frame: At the end of the canyon, the ground suddenly breaks into a tens-of-meters-wide abyss. The SUV charges at full speed up the natural rock ramp at the edge of the break and leaps into the air—slow motion: wheels spinning freely, dust separating and drifting off the body, sunlight passing through the chassis, background filled with an overwhelming sandstorm wall—the vehicle slams down heavily on the opposite side, suspension compressing violently, dust exploding outward, then continuing to speed away. [17–20 seconds: Dive push-in freeze · poster] Camera movement: Dive push-in + low-angle sudden stop, camera dives from high altitude back to the ground, skims over the sand, and abruptly stops at a low angle directly in front of the vehicle’s front. Frame: The SUV drives toward the camera and slowly brakes to a stop, kicked-up dust gradually settling. Frontal low-angle upward shot of the vehicle’s front, the paint surface reflecting the last trace of golden sunlight—the entire sky behind it has been completely swallowed by the sandstorm wall, the dark-golden storm like an apocalyptic sky curtain, only the car illuminated by sunlight. The frame freezes into a cinematic poster composition.
From a ladybug to wild horses
Fifteen seconds at sunset: a macro shot of a ladybug on grass that opens out to a herd running across the plain.
Show the prompt
Half an hour before sunset, the sky over the vast grassland gradually turned orange and purple. The grass swayed directionally with the actual wind speed (about level 3). The camera was positioned at a low position (about 20 cm off the ground), using a telephoto lens (200mm) to compress the space. The first five seconds begin with a close-up of a foxtail grass, its stem covered in fine hairs and adhering pollen grains. A ladybug crawls on it, its legs gripping the plant's intricate surface. In the background, distant mountains are hazy, clouds move slowly, and the light subtly shifts every second, lengthening the shadows of the grass blades. The next five seconds see the camera pan rapidly about 90 degrees, revealing a herd of wild horses (about five) trotting into the frame from the left. Their muscles undulate rhythmically with each stride, their manes flying as they run, each hair clearly defined. As their hooves trample the grass, they kick up dirt and grass clippings, creating Tyndall effect beams of dust in the air. Their breath condenses into white mist in the cool air. A foal follows its mother, its gait slightly unsteady, its ears darting to catch the wind. The final five seconds see the herd accelerate rapidly, the sense of speed conveyed through the rapid blurring of the foreground grass and the rapid rhythm of the hooves. The camera pulls back at approximately one meter per second. (out), gradually bringing the entire herd of horses into the frame, they rush toward the setting sun on the horizon. The backlight turns the horses into black silhouettes with gilded edges, their manes and tails fluttering in the wind. In the last two seconds, the sun completely sinks into the ridgeline, the light dims sharply, and the image fades into a deep purple-gray, leaving only a few distant neighs of horses.
Western standoff
Two gunfighters at dusk, dust in the wind and a hand hovering over a holster, framed like a widescreen classic.
Show the prompt
A desolate town at dusk, its main street shrouded in dust. Two cowboy figures stand facing each other at opposite ends of the street, about 50 meters apart. The shot is wide-angle and low-angle, emphasizing the vastness of the sky and the emptiness of the ground. The wind whips up dry grass and dust, creating swirling whirlwinds. A close-up shows one of the cowboys' hands hovering above his holster, his fingers trembling slightly. The setting sun casts long shadows of them onto the wooden buildings. Warm golden tones, backlit silhouettes, classic widescreen composition—the tension is frozen in the final second before the gun is drawn.
Aurora with its own score
Northern lights over a snowy forest, with a soft synth tone that the prompt asks to rise with the colors.
Show the prompt
A wide-angle, upward-looking shot captures the aurora borealis dancing slowly across the deep night sky like silk, its colors transitioning from emerald green to violet and then to rose pink, forming a flowing river of color. As the aurora crests gently rise, a warm, enveloping analog synthesizer tone resonates, its timbre soft and subtly tape-saturated, like a low sigh from the sky. As the aurora's colors subtly shift, the overtones of the sound also enrich, creating a perfect synesthetic experience. In the foreground are silhouettes of snow-covered coniferous forests, snowflakes shimmering in the soft light. The overall rhythm is extremely slow, like the deep breath of the universe, guiding the viewer into a deep alpha brainwave state. While the colors are rich, they are softened by light, not glaring, and filled with a sacred tranquility.
Golden-hour family moment
A father ties a ribbon round his daughter’s wrist at the kitchen table, in a slow push-in through window light.
Show the prompt
A 10-second cinematic shot, bathed in the iconic warmth of family intimacy. Golden-hour sunlight pours through the kitchen window, casting long amber beams across the wooden dining table. The camera slowly dollies forward at a child's eye level, gradually revealing a father kneeling on one knee beside his young daughter, their backlit silhouettes framed against the glowing window. He gently ties a small ribbon around her wrist; she looks up and smiles. Soft lens flares bloom, dust motes drifting lazily through the shafts of light. Shallow depth of field, anamorphic widescreen bokeh, the grain texture of Kodak Vision3 500T film. The camera continues its seamless, imperceptible push-in until their clasped hands fill the entire frame. Warm color grading dominated by honey gold and soft peach. Emotionally rich, tenderly intimate, timelessly nostalgic — like a memory you never realized you already had.
Shot for the feed, with the prompts
Generated at 9:16, not cropped from a wide frame.
Talking to camera
A UGC-style product review: she holds up a serum and talks about it, the voice generated with the picture.
Show the prompt
A woman in her late twenties sits in a bright bedroom holding a mint-green serum bottle with a gold cap up beside her face, talking straight to camera about it, turning the bottle to show the label. Vertical framing, handheld phone camera, soft daylight through window blinds behind her. Natural skin texture, a few loose strands of hair, an unforced smile between sentences. She speaks in a warm conversational voice.
Product unboxing
Hands open a matcha box and let the powder fall: packaging, texture and a quiet room tone.
Show the prompt
Two hands open a mint-green matcha latte box on a pale pink tabletop, lift the lid away, then dip fingers into a glass jar of matcha powder and let it fall back through in a fine stream. Vertical framing, overhead and slightly angled, soft diffused studio light. Powder clings to the fingertips, a little dusts the table, the jar glass is slightly smudged. Quiet room tone.
Ramen on the pass
Steam, chilli oil breaking the broth and the clink of ceramic — food footage with its own sound.
Show the prompt
A chef's hands finish a bowl of ramen on a small counter kitchen pass: broth steaming, chopsticks laying two slices of chashu, a ladle of chilli oil breaking the surface into red rings. Vertical framing, handheld 35mm, warm overhead tungsten against a cold window edge, shallow focus. Steam hazes the lens, a thumbprint smudges the bowl rim, the scallion scatter is uneven. Ambient kitchen room tone, the clink of ceramic on steel.
Golden-hour walk
A fashion shot: a camel coat in the wind on wet cobblestones, backlit, with street ambience.
Show the prompt
A woman in her late twenties walks toward camera down a narrow city street at golden hour, a long camel coat catching the wind, boots on wet cobblestone. Vertical framing, handheld 50mm follow, backlit with a low flare across the frame. Loose strands of hair cross her face, one boot toe is scuffed, the pavement reflects unevenly. Street ambience with distant traffic.
What is Wan AI, and what is Wan 3.0?
Wan AI is Alibaba’s family of video generation models, made by its Tongyi Lab (in Chinese, Tongyi Wanxiang). Wan 2.1 and Wan 2.2 came out as open weights in 2025 and became the go-to models for making video on your own GPU; the versions after them added native audio and longer takes. Alibaba runs Wan in its own app at wan.video and through Alibaba Cloud Model Studio.
Wan 3.0 is the newest release: a public beta from August 6, 2026, launched officially on August 24. It renders up to 30 seconds in one pass — twice Wan 2.7’s 15 — at up to 1080p, with sound generated alongside the picture, and it can follow reference images, clips, audio and even documents or web pages.
ToonBee AI runs Wan 3.0 through fal, so this is an unofficial way to use Wan AI: not Alibaba’s own service, and no Alibaba Cloud account needed. Here you can start from text, an image, a first and a last frame, or up to nine reference images, at 480p, 720p or 1080p and 4–30 seconds. Reference clips and audio, and the document and web-page inputs, are not offered here.
How to use Wan AI online
- 1
Write a prompt or add an image
Describe the scene, the action, the camera and the sound, and put anything a character says in quotes. Add an image to start from it, or set a first and a last frame.
- 2
Pick the size, length and shape
The box starts on Wan 3.0 at 480p and the shortest length, the cheapest clip. Choose 480p, 720p or 1080p and 4–30 seconds; the credit cost updates as you go.
- 3
Generate and download
Press Generate video. The clip lands in your history with its sound and downloads as an MP4 with no watermark.
Wan 3.0 vs the other video models here
Every video model you can pick on ToonBee AI, read from the generator’s own tables. Credits are for a 5-second clip, from the smallest size to the largest.
| Model | Max size | Length | Native audio | Image to video | First + last frame | Reference images | Credits (5s) |
|---|---|---|---|---|---|---|---|
| Wan 3.0 | 1080p | 4–30s | 125–500 | ||||
| Seedance 2.5 | 1080p | 5–30s | 238–1,328 | ||||
| H3 Max Turbo | 1080p | 5–15s | 63–200 | ||||
| Seedance 2 Mini | 1080p | 4–15s | 128–638 |
Wan 3.0 and Seedance 2.5 are the two models here that run to 30 seconds at 1080p with sound; Wan 3.0 does it for fewer credits. H3 Max Turbo is cheaper still, for clips of up to 15 seconds.
Wan 3.0 vs Wan 2.7
Wan 2.7 stops at 15 seconds a take. Wan 3.0 doubles that to 30 in one pass, with the sound made in the same generation. ToonBee AI runs Wan 3.0 only.
Wan 3.0 vs Seedance 2.5
Both make 30-second takes at 1080p with native audio. Opinions on the picture are split, but Wan 3.0 costs about half as many credits for the same clip here. Try one prompt on each from the same box.
Wan 3.0 vs H3 Max Turbo
H3 Max Turbo is the cheapest and quickest model on the site, up to 15 seconds a take. Wan 3.0 is for the longer take, up to 30.
Online vs running Wan locally
Wan 2.1 and 2.2 can run on your own GPU through ComfyUI. Wan 3.0 has no open weights, so it only runs as a service. Here it needs nothing installed.
Why use Wan AI here
Up to 30 seconds in one take
One continuous generation, not short clips stitched together, so camera moves and a whole scene hold together.
Native 1080p
Rendered at the size you pick — 480p, 720p or 1080p — with no upscale pass in between.
Sound in the same pass
Dialogue, effects and ambience are generated with the picture. Put a line in quotes and the character says it.
Text, image, frames or references
Start from a prompt, animate an image, pin a first and a last frame, or add up to nine reference images to keep a character or product consistent.
No Alibaba Cloud account
No cloud console, no API key, no GPU. It runs in any modern browser on a laptop, a PC or a phone.
The price before you press
The box shows the credit cost before you generate, and if a generation fails its credits come back automatically.
What people make with Wan AI
UGC-style ads
Someone talking to camera about a product, in 9:16, with the voice generated in the same take.
Product and food clips
Unboxings, close-ups and steaming plates with their own sound, without a shoot.
Short cinematic scenes
A full 30-second scene in one take — a chase, a reveal, a moment with dialogue.
Images brought to life
Start from a picture you made with an image model, or a photo, and set it in motion.
Vertical clips for the feed
Native 9:16 for TikTok, Reels and Shorts, generated upright rather than cropped.
Fashion and travel shots
Fabric in the wind, golden-hour light and street ambience, from one prompt.
Need a whole story with narration? Try the story video generator
Is Wan AI free? What does a clip cost?
Wan 3.0 has no free tier of its own: Alibaba and the platforms that carry it charge by the second. The older Wan 2.1 and 2.2 weights are free, but you need your own GPU to run them.
ToonBee AI is free to start: every new account gets 100 credits, with no card and no subscription — enough for a 4-second Wan 3.0 clip at 480p. A 5-second clip is 125 credits at 480p, 250 at 720p and 500 at 1080p; longer clips scale by the second, so a full 30-second take at 1080p is 3,000.
Plans start at $39 a month for 5,000 credits. Prefer not to subscribe? A one-time pack is $49 for 4,000 credits that never expire. Every plan covers commercial use and downloads with no watermark.
Good to know before you start
- One clip runs 4–30 seconds. For something longer, ToonBee 1.0 makes an edited video from 15 seconds, and the story video generator builds narrated videos of several minutes.
- Wan 3.0 runs with Alibaba’s content filter, which turns away some prompts and images — now and then harmless ones too. A refused generation gives its credits back.
- Reference input takes images only. Reference clips and audio, and Wan 3.0’s document and web-page inputs, are not offered here.
- Over a long take a character can drift. Start from an image, or add reference images, to keep a face consistent.
- Long 1080p takes take longer to render than short ones. Try the prompt at 480p first.
Wan AI FAQ
What is Wan AI?
Wan AI is a family of AI video generation models from Alibaba’s Tongyi Lab (Tongyi Wanxiang in Chinese). It turns text, images and other references into video clips. Earlier versions such as Wan 2.1 and 2.2 were released as open weights; the newest is Wan 3.0.
What is the latest Wan version?
Wan 3.0. It opened as a public beta on August 6, 2026 and launched on August 24, 2026. Before it, Wan 2.7 made clips of up to 15 seconds; Wan 3.0 makes up to 30 in one pass.
Is Wan AI free?
Wan 3.0 itself has no free tier; it is billed by the second wherever you use it. On ToonBee AI every new account gets 100 free credits with no card required, enough for a 4-second clip at 480p.
How much does a Wan 3.0 video cost?
Here a 5-second clip is 125 credits at 480p, 250 at 720p and 500 at 1080p, and the cost scales with length. Plans start at $39 a month for 5,000 credits.
How long can a Wan 3.0 video be, and what resolution?
Each clip is 4–30 seconds in one take, at 480p, 720p or 1080p, in 16:9, 4:3, 1:1, 3:4 or 9:16.
Does Wan 3.0 generate audio?
Yes. Sound is generated in the same pass as the picture: dialogue, effects and ambience. Write a line in quotes and the character says it.
Is Wan 3.0 open source? Can I run it locally?
Not so far. Wan 2.1 and 2.2 were released as open weights and can run on your own GPU, but Wan 3.0 is available only as a service, through Alibaba and platforms such as fal. Online, there is nothing to install.
Can Wan AI turn an image into a video?
Yes. Add an image and Wan 3.0 starts the clip from it; add a second one to set the last frame too. You can also add up to nine reference images for a character, product or place.
Wan 3.0 vs Seedance 2.5: which is better?
Both make 30-second takes at 1080p with sound, and people who compare them disagree on the picture. On ToonBee AI Wan 3.0 costs about half as many credits for the same clip, and both are in the same box, so you can run one prompt on each.
Why was my Wan prompt refused?
Wan 3.0 runs with Alibaba’s own content filter on prompts and images, and it can be strict. Rephrase the prompt or try another image. A refused generation gives its credits back.
Is this the official Wan AI site?
No. ToonBee AI is an independent video generator, not affiliated with Alibaba. It runs Wan 3.0 through fal. Alibaba’s own Wan app is at wan.video.
Can I use Wan AI videos commercially?
Yes. Videos you make on ToonBee AI can be used commercially on every plan, the free one included, and download with no watermark.
How do I write good Wan 3.0 prompts?
Write it in layers: who is in the shot, what they do, where it happens, the camera, the light, then the sound and any line in quotes. Concrete details — a scuffed boot, steam on the lens — help. The vertical samples on this page show full prompts.
Make your first Wan AI video
Type a prompt or add an image. Free credits to start, no GPU, no install.
Start with the free creditsRelated tools
Other pages that open the same studio, set up for a different kind of video.
- Dola AI Video GeneratorDola-style Seedance videos online, with no watermark.
- MiniMax H3 Video GeneratorMiniMax H3 videos with sound, on H3 Max in your browser.
- AI ModelsEvery video and image model, with lengths and prices.
- AI Animation GeneratorShort animated clips, or a still image brought to life.
- AI Movie MakerCinematic short films from a story idea, with narration.
- AI Cartoon Video GeneratorTurn an idea or a script into a narrated cartoon video.
ToonBee AI is an independent video generator. It runs Alibaba’s Wan 3.0 through fal. It is not the official Wan site (wan.video) and is not affiliated with or endorsed by Alibaba. Wan is a trademark of its owner.