Generated sound
Switch Audio on and the clip comes with ambience and effects that match the picture.
New
Muse’s flagship multimodal video model. Direct a scene in words, anchor it with frames or references, and get a clip with light, motion, and sound already composed.
Avoid copyrighted characters, logos, music and sensitive or explicit content — it may be blocked.
Estimated cost: 5 credits
Your creation appears here
Pick a model, describe the shot, and hit Generate.
Why Muse Video
Muse Video combines prompt, frame, and reference control with generated sound, inside one studio.
Switch Audio on and the clip comes with ambience and effects that match the picture.
Anchor a character, outfit, and art direction with a large reference set.
Draft at 480p, produce at 720p, and deliver at 1080p from the same prompt.
Start from text, a first frame, or references without changing tools.
Muse Video guide
Muse Video is the flagship video model in the Nano Banana AI studio. It accepts a text prompt, a first frame, or a set of reference stills, and it can generate sound together with the picture. This guide explains what each input does and when Muse Video is the right engine for a shot.
Muse Video covers three jobs in one model: text to video, image to video, and reference to video. You can describe a scene from scratch, animate a still you already have, or hand the model up to 30 reference images so a character and a look stay consistent across shots.
Output runs at 480p for cheap drafts, 720p for production work, and 1080p for final delivery. Aspect ratios include 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and adaptive, so one idea can be rendered for a cinema-style frame and a vertical feed.
Turn Audio on and the clip is generated with matching ambience and effects, so you are not syncing a track afterwards.
Up to 30 reference stills hold identity, wardrobe, and style across a sequence.
Test at 480p, then re-run the keeper at 720p or 1080p.
Write the shot the way a director would call it: who is in frame, what they do, how the camera moves, and what the light is doing. "A chef plates a dessert in a dim kitchen, slow push-in on the hands, warm tungsten light, steam rising" is specific enough to direct the clip and short enough to stay coherent.
When sound is on, name it. Rain on a window, a crowd murmur, or a quiet room tone gives the audio generator something to match instead of guessing.
Open the Reference tab and add clean stills: one clear face, one outfit, one location. Say what each is for in the prompt. Three well-chosen references beat thirty near-identical portraits, and consistent lighting across the references makes the result more stable.
Pick Muse Video when you need audio, many references, or 1080p. Pick Seedance 2.5 for long clips, Veo 3.1 for quick realistic short shots, Wan 3.0 for stylized work, and MiniMax H3 for fluid motion at 2K. Everything lives in the same studio, so you can try one idea on two models and keep the better take.
How to use Muse Video
State the subject, action, camera move, and light. Add the sounds you want if Audio is on.
Who uses Muse Video

Previs, mood reels, and establishing shots with sound already sketched in.

Vertical and widescreen versions of one concept, generated from the same references.

Product and lifestyle clips that keep packaging and palette consistent.

Recurring characters across scenes using a shared reference set.
More models
Switch engines without switching tools — Nano Banana 2.1 plus the strongest public image and video models, side by side in the Nano Banana AI studio.
FAQ