wan ·

    Wan 3.0 Is Coming: 30 Seconds at 1080p, Sound in the Same Pass

    Alibaba's next-generation video model is in beta now. One continuous 30-second take at 1080p, audio generated with the picture, every ratio from 16:9 to 9:16 - and a thinking mode that plans the shot before rendering it.

    Adam Balogh4 min read
    Wan 3.0 Is Coming: 30 Seconds at 1080p, Sound in the Same Pass

    Alibaba's Tongyi Lab - the team behind every Wan release from 2.1 through 2.7 - has opened the beta for Wan 3.0. It isn't publicly available yet: right now it's a closed beta on Alibaba's own platforms, with wider availability expected to follow. But the published specs are enough to say this is the release worth planning around, so here's what's coming and where it'll land when it does.

    One take, thirty seconds, no stitching

    The headline number: anywhere from 2 to 30 seconds in a single generation. Not short clips glued together afterwards - one continuous pass. That's double Wan 2.7's 15-second ceiling, and matching the longest single take of any major model right now.

    The length itself isn't really the point, though. One continuous pass is what lets camera language survive a whole scene. A push-in that keeps pushing, a tracking shot that actually tracks, a one-take that holds its nerve for the full thirty. When a scene is stitched from 5-second fragments, that grammar dies at every seam. When it's one pass, it lives.

    There's also a small feature I suspect people will end up loving: leave the duration unset and the model reads your prompt and picks its own. A two-line gag gets four seconds, a scene brief gets what a scene needs.

    1080p with the sound generated alongside the picture

    Output runs at 480p, 720p, or 1080p, with 1080p as the default - and audio is generated in the same pass as the video. Dialogue, ambience, and on-screen action land together instead of being dubbed on afterwards. Wan 2.7 already does native audio well; 3.0 extends it across the longer window with voices that stay consistent through the take.

    Aspect ratios cover 16:9, 4:3, 1:1, 3:4, and 9:16 - native portrait generation, not a landscape frame cropped down - or you can let the model pick the ratio that suits the shot.

    It thinks before it renders

    The most interesting line in the spec sheet: Wan 3.0 has a thinking mode. Turn it on and the model reasons about composition and motion before it renders a frame. Alibaba's demo material leads with the hardest test there is for this - fast, athletic, full-body movement, a night dance battle in a crowd under mixed lighting - precisely the kind of shot where current models turn limbs into soup. Planning the motion before committing pixels is how you keep weight, contact, and momentum readable through a whole take.

    Alibaba's beta notes go further still: alongside text, images, audio, and video, Wan 3.0 reportedly accepts documents - PDFs, decks, spreadsheets - as reference input, turning static material directly into video. We'll believe the details when we can run them, but the ambition is clear.

    The open question - literally

    The thing we care most about hasn't been answered: whether Wan 3.0 gets an open release. The Wan family earned its place as the workhorse of uncensored AI video because open weights mean providers can host it without a moderation layer second-guessing your prompts. Wan 3.0 is API-only beta for now, and Alibaba hasn't said what happens after.

    Until that's confirmed, assume the unrestricted side of the house stays on Wan 2.7 - which remains excellent, fully live in our Video Studio, and covered in our no-code guide.

    What happens when it ships

    Same as every model drop: we bring it up as soon as API access opens, it shows up in the Video Studio picker, and the price in credits shows before you render (1,000 credits = $1). And as always, it runs behind our anonymity layer - encrypted on your device, identity split from content by Oblivious HTTP relays, decrypted only inside attested secure enclaves. The full architecture is here. A model that can hold a 30-second one-take is a model people will bring real work to. That work shouldn't be logged against your name.

    While you wait

    The best preparation for Wan 3.0 is getting good at writing scenes rather than prompts - the skill transfers directly from the current generation:

    Try Wan 2.7 text to video

    Read the AI video prompt guide

    We'll publish again the day Wan 3.0 goes live here.

    • #wan
    • #alibaba
    • #video
    • #announcement
    • #video-studio

    Keep reading