Wan AI: Generador de Videos con Wan 3.0 en Línea y con Sonido

Crea videos con Wan AI (Wan IA) desde tu navegador con Wan 3.0, el modelo Wan más reciente de Alibaba. Escribe un prompt o añade una imagen y obtén una sola toma continua de hasta 30 segundos y hasta 1080p, con diálogos, efectos y ambiente generados a la vez. Sin cuenta de Alibaba Cloud, sin GPU y sin instalar nada.

  • 100 créditos gratis
  • Hasta 30s en una toma
  • 1080p nativo con sonido
  • De texto o imagen a video

El contenido sexualmente explícito (NSFW) está estrictamente prohibido.

Ejemplos de Wan 3.0

Clips hechos con Wan 3.0, con su sonido y con el prompt que los creó. Pasa el cursor para reproducir, activa el sonido, abre el prompt o ponlo en la caja y crea tu propia versión.

Tomas únicas largas, de hasta 30 segundos

Cada una es una sola generación, no planos montados.

30sPasa el cursor o toca para reproducir

30 segundos en una toma

El vuelo de un águila desde un acantilado nevado, entre un tigre dientes de sable y dos alces gigantes en pelea, hasta un volcán en erupción, sin un solo corte.

Ver el prompt

Global Style & Setting: Photorealistic epic nature documentary style (like Prehistoric Planet or Planet Earth). 30 seconds of CG epic one-shot. We follow the FPV perspective of a Haast's Eagle (the largest eagle in history, with a 3-meter wingspan), flying continuously from a snowy cliff nest, across the Pleistocene glacial forests, interacting thrillingly with prehistoric megafauna. The camera blends high-speed flight, elegant gliding, thrilling dodges, and grand reveals. Lighting transitions from the warm golden hour of sunset to the cold blue of glacial twilight, ending with the fiery red of a volcanic eruption. [00.00-07.00s] Departure & Gliding: The Snowy Cliff Prompt: [FPV POV, One continuous shot]. The camera starts in a massive nest made of rocks and dry branches on a snowy cliff edge. Several fluffy Haast's eagle chicks are screeching beside us. With a high-pitched cry, we (the eagle) spread our massive, feathered wings, take a few running steps, and leap off the cliff. The camera dives into an extreme, heart-pounding freefall. After a moment of weightlessness, we spread our wings, catching the updraft to glide smoothly. The camera transitions to a steady, forward-tracking climb. Below us is the turbulent, icy blue ocean, with the setting sun casting shimmering golden light on the waves. Details: We can see the primary feathers trembling in the wind. The sound of howling wind mixes with the crashing of waves against the ice. [07.00-14.00s] Skimming the Ice & The Ambush Prompt: We drop altitude, flying almost skimming the surface of a frozen glacial lake. The camera arcs right, counter-clockwise, tracking the shoreline. On the ice, a herd of giant, flightless Moas (terror birds) are foraging. Suddenly, a massive shadow bursts from the snowdrifts—a Smilodon (saber-toothed cat) leaps out, ambushing a lone Moa. A massive explosion of snow and ice shards erupts, almost splashing the lens. We react instantly, pulling up sharply in an emergency ascent, narrowly dodging the flying ice. Lighting/Performance: The Moas scatter in panic. The Smilodon sinks back into the snow. The setting sun refracts through the suspended ice mist, creating a brief, stunning ice-rainbow. [14.00-23.00s] Forest Chase & The Antler Gap Escape Prompt: Flying over the ice, we dive into a dense, lush forest of prehistoric Araucaria trees. To dodge a rival, more aggressive Haast's Eagle patrolling the canopy, we plunge beneath the tree line. The camera switches to extreme high-speed FPV drone mode, weaving rapidly between giant ferns and towering trunks. Suddenly, we burst from behind a tree into a clearing. Two colossal Megaloceros (Irish Elks with 12-foot antlers) are locking horns in a violent, earth-shaking battle. [The Antler Gap穿越] We try to fly around, but the elks charge directly into our path! Time seems to slow down. The camera executes a dolly zoom (Vertigo effect); the forest background rushes backward, while the massive, interlocking, tree-branch-like antlers seem to fill the entire world, crushing toward us! With nowhere to go, we choose the only path—flying directly through the narrow, shifting gap between their crashing antlers! Details: The camera猛地 (suddenly) dives into the gap. For 0.5 seconds, the frame is surrounded by the rough, moss-covered texture of the giant antlers, snapping twigs, and flying snow. Just milliseconds before the antlers crush together with a deafening "CRACK," we slip out through the bottom gap! The camera shakes violently from the shockwave, and a few snowflakes stick to the lens. [23.00-30.00s] Soaring Over Glaciers & The Epic Panorama Prompt: Escaping by a hair's breadth, we pull up frantically, gaining altitude to leave the danger below. The Megaloceros let out a furious, bellowing roar. Riding the thermal currents, we begin a spiral climb. As our altitude increases, the entire glacial archipelago is revealed. The camera pulls up to reveal a distant volcano erupting, with lava flowing like red rivers down the black rock. Below on the frozen plains, thousands of Moas are migrating in a massive herd. Finally, exhausted but reborn, we fly into the dark red clouds stained by volcanic ash and the glowing Aurora Borealis, leaving only a silhouette. The screen fades to black. Lighting/Emotion: The entire world is dyed in a magnificent, apocalyptic red by the sunset and the volcano. The sound of our heavy, post-survival breathing mixes with the wind, adding a sense of life's fragility and resilience to the grand finale.

20sPasa el cursor o toca para reproducir

Persecución en el desierto, 20 segundos

Un dron sigue a un todoterreno por un cañón de roca roja mientras una tormenta de arena se traga las ruinas detrás.

Ver el prompt

[One-take, FPV drone chase tracking shot, no cuts, 20-second long take, post-apocalyptic wasteland cinematic feel] Overall parameters: 20 seconds, 4K 60fps, 16:9 landscape. Realistic post-apocalyptic wasteland style; epic lighting, physics-level dust particle effects, realistic vehicle dynamics and suspension details, warm golden wasteland tones, shallow depth of field, extreme car body paint and material details. Global setting: A post-apocalyptic red rock canyon, with weathered ancient cave ruins on the cliff walls on both sides, low-angle golden afternoon sunlight. On the distant horizon, a dark-golden sandstorm wall several thousand meters high is advancing, swallowing everything in its path. The subject is a silver-white futuristic luxury SUV, streamlined body, closed grille, no brand logos or text anywhere on the vehicle, body covered with a thin layer of dust, driverless, moving autonomously. [0–3 seconds: Wheel close-up start] Camera movement: Extremely low-angle close-up, camera close to the ground aimed at the wheels. Frame: The SUV’s wheels roll over the sandy ground, slowly starting and accelerating, dust kicked up from the tires, with the mechanical details of the suspension and wheel hubs clearly visible. In the distant background behind the wheels, the sandstorm wall presses down on the horizon like a sky curtain, advancing toward the canyon. [3–7 seconds: Low-altitude side tracking shot · chase] Camera movement: Low-altitude parallel side tracking, camera moving sideways at the same speed as the vehicle. Frame: The SUV drives at high speed through the canyon, paint surface shimmering with reflected light, the rear dragging a long golden dust trail. The sandstorm wall is visibly closer—it swallows the cliff walls and cave ruins it passes, ancient caves disappearing one by one in the sandstorm, the unswallowed cliff walls on both sides receding rapidly, maximizing the oppressive feeling of the approaching storm. [7–10 seconds: Through-window shot · eerie] Camera movement: Fly-through shot, camera accelerates to catch up with the vehicle, smoothly passes through the rear window into the cockpit, sweeps across the interior, then exits through the windshield. Frame: The camera passes through the cabin—an empty driverless cockpit, wraparound ambient lighting and the horizontal large screen faintly glowing, leather and metal texture details visible; the instant it exits through the windshield, the sandstorm wall has already filled the entire far end of the forward view, with a faint rumbling sound coming from afar. [10–14 seconds: Drift orbit · giant wreckage] Camera movement: Orbit shot, the SUV drifts with its rear swinging out, camera orbits 180 degrees around the vehicle as the center. Frame: Ahead on the sandy ground, a hundreds-of-meters-tall giant mechanical wreckage is diagonally embedded (a half-weathered mechanical arm, covered in rust and dust). The SUV drifts around the wreckage, body sliding sideways, dust violently rising in a ring and covering half the frame, sunlight piercing through the sand curtain forming a golden halo, the shadow of the wreckage and the sandstorm wall overlapping in the background. [14–17 seconds: Cliff leap · slow motion] Camera movement: Vertical rise followed by slow motion, camera lifts with the vehicle as it becomes airborne. Frame: At the end of the canyon, the ground suddenly breaks into a tens-of-meters-wide abyss. The SUV charges at full speed up the natural rock ramp at the edge of the break and leaps into the air—slow motion: wheels spinning freely, dust separating and drifting off the body, sunlight passing through the chassis, background filled with an overwhelming sandstorm wall—the vehicle slams down heavily on the opposite side, suspension compressing violently, dust exploding outward, then continuing to speed away. [17–20 seconds: Dive push-in freeze · poster] Camera movement: Dive push-in + low-angle sudden stop, camera dives from high altitude back to the ground, skims over the sand, and abruptly stops at a low angle directly in front of the vehicle’s front. Frame: The SUV drives toward the camera and slowly brakes to a stop, kicked-up dust gradually settling. Frontal low-angle upward shot of the vehicle’s front, the paint surface reflecting the last trace of golden sunlight—the entire sky behind it has been completely swallowed by the sandstorm wall, the dark-golden storm like an apocalyptic sky curtain, only the car illuminated by sunlight. The frame freezes into a cinematic poster composition.

15sPasa el cursor o toca para reproducir

De una mariquita a caballos salvajes

Quince segundos al atardecer: un plano macro de una mariquita en la hierba que se abre a una manada corriendo por la llanura.

Ver el prompt

Half an hour before sunset, the sky over the vast grassland gradually turned orange and purple. The grass swayed directionally with the actual wind speed (about level 3). The camera was positioned at a low position (about 20 cm off the ground), using a telephoto lens (200mm) to compress the space. The first five seconds begin with a close-up of a foxtail grass, its stem covered in fine hairs and adhering pollen grains. A ladybug crawls on it, its legs gripping the plant's intricate surface. In the background, distant mountains are hazy, clouds move slowly, and the light subtly shifts every second, lengthening the shadows of the grass blades. The next five seconds see the camera pan rapidly about 90 degrees, revealing a herd of wild horses (about five) trotting into the frame from the left. Their muscles undulate rhythmically with each stride, their manes flying as they run, each hair clearly defined. As their hooves trample the grass, they kick up dirt and grass clippings, creating Tyndall effect beams of dust in the air. Their breath condenses into white mist in the cool air. A foal follows its mother, its gait slightly unsteady, its ears darting to catch the wind. The final five seconds see the herd accelerate rapidly, the sense of speed conveyed through the rapid blurring of the foreground grass and the rapid rhythm of the hooves. The camera pulls back at approximately one meter per second. (out), gradually bringing the entire herd of horses into the frame, they rush toward the setting sun on the horizon. The backlight turns the horses into black silhouettes with gilded edges, their manes and tails fluttering in the wind. In the last two seconds, the sun completely sinks into the ridgeline, the light dims sharply, and the image fades into a deep purple-gray, leaving only a few distant neighs of horses.

10sPasa el cursor o toca para reproducir

Duelo del oeste

Dos pistoleros al anochecer, polvo al viento y una mano sobre la funda, encuadrados como un clásico en pantalla ancha.

Ver el prompt

A desolate town at dusk, its main street shrouded in dust. Two cowboy figures stand facing each other at opposite ends of the street, about 50 meters apart. The shot is wide-angle and low-angle, emphasizing the vastness of the sky and the emptiness of the ground. The wind whips up dry grass and dust, creating swirling whirlwinds. A close-up shows one of the cowboys' hands hovering above his holster, his fingers trembling slightly. The setting sun casts long shadows of them onto the wooden buildings. Warm golden tones, backlit silhouettes, classic widescreen composition—the tension is frozen in the final second before the gun is drawn.

10sPasa el cursor o toca para reproducir

Aurora con su propia música

Auroras boreales sobre un bosque nevado, con un tono suave de sintetizador que el prompt pide que crezca con los colores.

Ver el prompt

A wide-angle, upward-looking shot captures the aurora borealis dancing slowly across the deep night sky like silk, its colors transitioning from emerald green to violet and then to rose pink, forming a flowing river of color. As the aurora crests gently rise, a warm, enveloping analog synthesizer tone resonates, its timbre soft and subtly tape-saturated, like a low sigh from the sky. As the aurora's colors subtly shift, the overtones of the sound also enrich, creating a perfect synesthetic experience. In the foreground are silhouettes of snow-covered coniferous forests, snowflakes shimmering in the soft light. The overall rhythm is extremely slow, like the deep breath of the universe, guiding the viewer into a deep alpha brainwave state. While the colors are rich, they are softened by light, not glaring, and filled with a sacred tranquility.

10sPasa el cursor o toca para reproducir

Un momento familiar a la hora dorada

Un padre le ata una cinta en la muñeca a su hija en la mesa de la cocina, con un travelling lento a contraluz.

Ver el prompt

A 10-second cinematic shot, bathed in the iconic warmth of family intimacy. Golden-hour sunlight pours through the kitchen window, casting long amber beams across the wooden dining table. The camera slowly dollies forward at a child's eye level, gradually revealing a father kneeling on one knee beside his young daughter, their backlit silhouettes framed against the glowing window. He gently ties a small ribbon around her wrist; she looks up and smiles. Soft lens flares bloom, dust motes drifting lazily through the shafts of light. Shallow depth of field, anamorphic widescreen bokeh, the grain texture of Kodak Vision3 500T film. The camera continues its seamless, imperceptible push-in until their clasped hands fill the entire frame. Warm color grading dominated by honey gold and soft peach. Emotionally rich, tenderly intimate, timelessly nostalgic — like a memory you never realized you already had.

En vertical para redes, con sus prompts

Generados en 9:16, no recortados de un plano ancho.

8sPasa el cursor o toca para reproducir

Hablando a cámara

Una reseña de producto estilo UGC: muestra un sérum y habla de él, con la voz generada junto a la imagen.

Ver el prompt

A woman in her late twenties sits in a bright bedroom holding a mint-green serum bottle with a gold cap up beside her face, talking straight to camera about it, turning the bottle to show the label. Vertical framing, handheld phone camera, soft daylight through window blinds behind her. Natural skin texture, a few loose strands of hair, an unforced smile between sentences. She speaks in a warm conversational voice.

6sPasa el cursor o toca para reproducir

Unboxing de producto

Unas manos abren una caja de matcha y dejan caer el polvo: envase, textura y un ambiente silencioso.

Ver el prompt

Two hands open a mint-green matcha latte box on a pale pink tabletop, lift the lid away, then dip fingers into a glass jar of matcha powder and let it fall back through in a fine stream. Vertical framing, overhead and slightly angled, soft diffused studio light. Powder clings to the fingertips, a little dusts the table, the jar glass is slightly smudged. Quiet room tone.

6sPasa el cursor o toca para reproducir

Ramen recién servido

Vapor, aceite de chile sobre el caldo y el tintineo de la cerámica: imágenes de comida con su propio sonido.

Ver el prompt

A chef's hands finish a bowl of ramen on a small counter kitchen pass: broth steaming, chopsticks laying two slices of chashu, a ladle of chilli oil breaking the surface into red rings. Vertical framing, handheld 35mm, warm overhead tungsten against a cold window edge, shallow focus. Steam hazes the lens, a thumbprint smudges the bowl rim, the scallion scatter is uneven. Ambient kitchen room tone, the clink of ceramic on steel.

6sPasa el cursor o toca para reproducir

Paseo a la hora dorada

Un plano de moda: un abrigo camel al viento sobre adoquines mojados, a contraluz y con sonido de calle.

Ver el prompt

A woman in her late twenties walks toward camera down a narrow city street at golden hour, a long camel coat catching the wind, boots on wet cobblestone. Vertical framing, handheld 50mm follow, backlit with a low flare across the frame. Loose strands of hair cross her face, one boot toe is scuffed, the pavement reflects unevenly. Street ambience with distant traffic.

¿Qué es Wan AI y qué es Wan 3.0?

Wan AI (también escrito Wan IA) es la familia de modelos de generación de video de Alibaba, creada por su laboratorio Tongyi (en chino, Tongyi Wanxiang). Wan 2.1 y Wan 2.2 se publicaron con pesos abiertos en 2025 y se convirtieron en los modelos favoritos para crear video con la GPU propia; las versiones siguientes añadieron audio nativo y tomas más largas. Alibaba ofrece Wan en su propia app, wan.video, y en Alibaba Cloud Model Studio.

Wan 3.0 es la versión más reciente: beta pública desde el 6 de agosto de 2026 y lanzamiento oficial el 24 de agosto. Genera hasta 30 segundos en una sola pasada —el doble que los 15 de Wan 2.7— y hasta 1080p, con el sonido generado junto a la imagen, y puede seguir imágenes, clips y audios de referencia, e incluso documentos o páginas web.

ToonBee AI ejecuta Wan 3.0 a través de fal, así que es una forma no oficial de usar Wan AI: no es el servicio de Alibaba y no necesitas cuenta de Alibaba Cloud. Aquí puedes partir de texto, una imagen, un primer y un último fotograma o hasta nueve imágenes de referencia, en 480p, 720p o 1080p y de 4 a 30 segundos. Los clips y audios de referencia, y la entrada de documentos y páginas web, no están disponibles aquí.

Cómo usar Wan AI en línea

  1. 1

    Escribe un prompt o añade una imagen

    Describe la escena, la acción, la cámara y el sonido, y pon entre comillas lo que diga un personaje. Añade una imagen para empezar desde ella, o fija un primer y un último fotograma.

  2. 2

    Elige tamaño, duración y formato

    La caja empieza con Wan 3.0 en 480p y la duración más corta, el clip más barato. Elige 480p, 720p o 1080p y de 4 a 30 segundos; el costo en créditos se actualiza al momento.

  3. 3

    Genera y descarga

    Pulsa Generar video. El clip aparece en tu historial con su sonido y se descarga en MP4 sin marca de agua.

Wan 3.0 frente a los otros modelos de video

Todos los modelos de video que puedes elegir en ToonBee AI, según las tablas del propio generador. Los créditos son para un clip de 5 segundos, del tamaño más pequeño al más grande.

ModeloTamaño máx.DuraciónAudio nativoImagen a videoPrimer + último fotogramaImágenes de referenciaCréditos (5s)
Wan 3.01080p4–30s125–500
Seedance 2.51080p5–30s238–1,328
H3 Max Turbo1080p5–15s63–200
Seedance 2 Mini1080p4–15s128–638

Wan 3.0 y Seedance 2.5 son los dos modelos de aquí que llegan a 30 segundos en 1080p con sonido; Wan 3.0 lo hace con menos créditos. H3 Max Turbo es aún más barato, para clips de hasta 15 segundos.

Wan 3.0 frente a Wan 2.7

Wan 2.7 se queda en 15 segundos por toma. Wan 3.0 lo duplica a 30 en una sola pasada, con el sonido generado a la vez. ToonBee AI solo ejecuta Wan 3.0.

Wan 3.0 frente a Seedance 2.5

Los dos hacen tomas de 30 segundos en 1080p con audio nativo. Las opiniones sobre la imagen están divididas, pero aquí Wan 3.0 cuesta más o menos la mitad de créditos por el mismo clip. Prueba un prompt en cada uno desde la misma caja.

Wan 3.0 frente a H3 Max Turbo

H3 Max Turbo es el modelo más barato y rápido del sitio, hasta 15 segundos por toma. Wan 3.0 es para la toma más larga, de hasta 30.

En línea o Wan en local

Wan 2.1 y 2.2 pueden ejecutarse en tu propia GPU con ComfyUI. Wan 3.0 no tiene pesos abiertos, así que solo funciona como servicio. Aquí no hay que instalar nada.

Compara todos los modelos del sitio

Por qué usar Wan AI aquí

Hasta 30 segundos en una toma

Una sola generación continua, no clips cortos unidos, así que los movimientos de cámara y la escena entera se mantienen.

1080p nativo

Se renderiza al tamaño que eliges —480p, 720p o 1080p— sin reescalado intermedio.

Sonido en la misma pasada

Diálogos, efectos y ambiente se generan con la imagen. Pon una frase entre comillas y el personaje la dice.

Texto, imagen, fotogramas o referencias

Parte de un prompt, anima una imagen, fija un primer y un último fotograma o añade hasta nueve imágenes de referencia para mantener un personaje o un producto.

Sin cuenta de Alibaba Cloud

Sin consola en la nube, sin clave de API y sin GPU. Funciona en cualquier navegador moderno, en portátil, PC o móvil.

El precio antes de pulsar

La caja muestra el costo en créditos antes de generar, y si una generación falla, los créditos se devuelven solos.

Qué se crea con Wan AI

Anuncios estilo UGC

Alguien que habla a cámara sobre un producto, en 9:16, con la voz generada en la misma toma.

Clips de producto y comida

Unboxings, primeros planos y platos humeantes con su propio sonido, sin rodaje.

Escenas cinematográficas cortas

Una escena completa de 30 segundos en una toma: una persecución, una revelación, un momento con diálogo.

Imágenes que cobran vida

Parte de una imagen que creaste con un modelo de imagen, o de una foto, y ponla en movimiento.

Clips verticales para redes

9:16 nativo para TikTok, Reels y Shorts, generado en vertical y no recortado.

Moda y viajes

Tela al viento, luz de hora dorada y sonido de calle, a partir de un solo prompt.

¿Necesitas una historia completa con narración? Prueba el generador de videos de historias

¿Wan AI es gratis? ¿Cuánto cuesta un clip?

Wan 3.0 no tiene un plan gratuito propio: Alibaba y las plataformas que lo ofrecen cobran por segundo. Los pesos de Wan 2.1 y 2.2 son gratis, pero necesitas tu propia GPU para usarlos.

ToonBee AI es gratis para empezar: cada cuenta nueva recibe 100 créditos, sin tarjeta ni suscripción, suficientes para un clip de 4 segundos con Wan 3.0 en 480p. Un clip de 5 segundos cuesta 125 créditos en 480p, 250 en 720p y 500 en 1080p; los clips más largos se cobran por segundo, así que una toma completa de 30 segundos en 1080p cuesta 3000.

Los planes empiezan en $39 al mes por 5000 créditos. ¿Prefieres no suscribirte? Un paquete único cuesta $49 por 4000 créditos que no caducan. Todos los planes permiten uso comercial y descargas sin marca de agua.

Ver todos los planes

Antes de empezar

  • Cada clip dura de 4 a 30 segundos. Para algo más largo, ToonBee 1.0 crea un video editado desde 15 segundos, y el generador de videos de historias arma videos narrados de varios minutos.
  • Wan 3.0 funciona con el filtro de contenido de Alibaba, que rechaza algunos prompts e imágenes, y a veces también algunos inofensivos. Una generación rechazada devuelve sus créditos.
  • La entrada de referencia solo admite imágenes. Los clips y audios de referencia, y la entrada de documentos y páginas web de Wan 3.0, no están disponibles aquí.
  • En una toma larga un personaje puede cambiar. Parte de una imagen o añade imágenes de referencia para mantener la cara.
  • Las tomas largas en 1080p tardan más en renderizarse que las cortas. Prueba primero el prompt en 480p.

Preguntas frecuentes sobre Wan AI

¿Qué es Wan AI (Wan IA)?

Wan AI es una familia de modelos de generación de video con IA del laboratorio Tongyi de Alibaba (Tongyi Wanxiang en chino). Convierte texto, imágenes y otras referencias en clips de video. Versiones anteriores como Wan 2.1 y 2.2 se publicaron con pesos abiertos; la más reciente es Wan 3.0.

¿Cuál es la última versión de Wan?

Wan 3.0. Abrió como beta pública el 6 de agosto de 2026 y se lanzó el 24 de agosto de 2026. Antes, Wan 2.7 hacía clips de hasta 15 segundos; Wan 3.0 llega a 30 en una sola pasada.

¿Wan AI es gratis?

Wan 3.0 no tiene plan gratuito propio; se cobra por segundo en cualquier plataforma. En ToonBee AI cada cuenta nueva recibe 100 créditos gratis sin tarjeta, suficientes para un clip de 4 segundos en 480p.

¿Cuánto cuesta un video con Wan 3.0?

Aquí un clip de 5 segundos cuesta 125 créditos en 480p, 250 en 720p y 500 en 1080p, y el costo crece con la duración. Los planes empiezan en $39 al mes por 5000 créditos.

¿Cuánto puede durar un video de Wan 3.0 y en qué resolución?

Cada clip dura de 4 a 30 segundos en una toma, en 480p, 720p o 1080p, y en 16:9, 4:3, 1:1, 3:4 o 9:16.

¿Wan 3.0 genera audio?

Sí. El sonido se genera en la misma pasada que la imagen: diálogos, efectos y ambiente. Escribe una frase entre comillas y el personaje la dice.

¿Wan 3.0 es de código abierto? ¿Puedo usarlo en local?

Por ahora no. Wan 2.1 y 2.2 se publicaron con pesos abiertos y pueden ejecutarse en tu propia GPU, pero Wan 3.0 solo está disponible como servicio, a través de Alibaba y de plataformas como fal. En línea no hay nada que instalar.

¿Wan AI puede convertir una imagen en video?

Sí. Añade una imagen y Wan 3.0 empieza el clip desde ella; añade otra para fijar también el último fotograma. También puedes añadir hasta nueve imágenes de referencia de un personaje, un producto o un lugar.

Wan 3.0 o Seedance 2.5: ¿cuál es mejor?

Los dos hacen tomas de 30 segundos en 1080p con sonido, y quienes los comparan no se ponen de acuerdo sobre la imagen. En ToonBee AI Wan 3.0 cuesta más o menos la mitad de créditos por el mismo clip, y los dos están en la misma caja, así que puedes probar un prompt en cada uno.

¿Por qué se rechazó mi prompt en Wan?

Wan 3.0 aplica el filtro de contenido propio de Alibaba a prompts e imágenes, y puede ser estricto. Reformula el prompt o prueba otra imagen. Una generación rechazada devuelve sus créditos.

¿Es este el sitio oficial de Wan AI?

No. ToonBee AI es un generador de videos independiente, sin relación con Alibaba. Ejecuta Wan 3.0 a través de fal. La app oficial de Wan está en wan.video.

¿Puedo usar los videos de Wan AI con fines comerciales?

Sí. Los videos que creas en ToonBee AI se pueden usar comercialmente en todos los planes, incluido el gratuito, y se descargan sin marca de agua.

¿Cómo escribo buenos prompts para Wan 3.0?

Escríbelo por capas: quién sale en el plano, qué hace, dónde ocurre, la cámara, la luz y luego el sonido y cualquier frase entre comillas. Los detalles concretos —una bota rayada, vapor en la lente— ayudan. Los ejemplos verticales de esta página muestran los prompts completos.

Crea tu primer video con Wan AI

Escribe un prompt o añade una imagen. Créditos gratis para empezar, sin GPU ni instalación.

Empezar con los créditos gratis

ToonBee AI es un generador de videos independiente. Usa Wan 3.0 de Alibaba a través de fal. No es el sitio oficial de Wan (wan.video) y no está afiliado ni respaldado por Alibaba. Wan es una marca de su propietario.