Veo 3.1: Google's Video Model with Sound

Veo 3.1 is Google's video generation model. It produces photoreal shots of up to twelve seconds, and unlike every other model here it generates sound along with the picture. Run it inside MarsClip and the shot lands in a project that already handles script, voiceover, captions and publishing.

How to generate a video with Veo 3.1

  1. Start from a script or a topic

    A shot is easier to describe when you know what it has to say. Bring a script, or have one written from a topic, and the video is broken into scenes before anything is generated. Each scene becomes a prompt, which is a far better starting point than a blank box and a hope.

  2. Choose Veo 3.1 for the scene

    The model is picked per scene rather than per project, so a video can open on a Veo shot with its own ambience and continue on another model where length or a repeating character matters more. Nothing forces one choice across the whole piece.

  3. Describe the shot, and say what it sounds like

    Write what is in frame, where the camera is and what is moving. Then describe the audio, because Veo is generating that too. Rain on a metal roof, a room tone, a distant road: naming the sound is the part people forget, and it is the reason to be using this model at all.

  4. Assemble, narrate and caption

    Generated shots are laid out against the script. Voiceover is generated in any of 32 languages, captions are timed word by word to that voiceover, and the finished video is rendered with Remotion on AWS Lambda rather than on your machine.

  5. Publish on a schedule

    Connect YouTube, TikTok and Instagram and set when the video goes live. Series and automation can keep a channel posting without you opening the app, which is usually the step that decides whether a channel lasts past its third week.

What Veo 3.1 is good for

Shots where the sound is the point

A wave breaking, a kettle, footsteps in a stairwell, a market at midday. When a scene is carried by its atmosphere rather than by narration, generating picture and audio together saves you sourcing a sound effect that never quite matches the picture you ended up with.

Photoreal establishing shots

The opening frame of a piece has to convince quickly. Veo leans photoreal, which suits the wide shot that sets a place before a narrator starts talking over it.

B-roll for narration-led video

Most faceless video is a script over pictures. Veo shots slot under a voiceover cleanly. Where a moment wants its own diegetic sound, you have it without leaving the project: traffic, weather, a machine.

Product and brand moments

A short, well-lit shot of a thing behaving the way it should. Twelve seconds is more than enough for the single beat most brand clips are built from, and the accompanying audio gives it presence that a silent generated clip lacks.

Ads and social spots

Short-form rewards a strong first second. A photoreal shot with its own sound gives you something to open on that does not read as stock, and the vertical cut, captions and scheduled post are handled in the same project.

Testing a look before committing

Generating one Veo shot is a cheap way to find out whether a scene works at all. If the idea does not survive twelve seconds, it will not survive a whole video, and you have learned that before writing the rest.

What is Veo 3.1?

Veo 3.1 is a video generation model made by Google. You give it a written description of a shot, or an image to start from, and it returns a short video clip. Within MarsClip it generates clips of up to twelve seconds.

What separates it from the rest of the models available here is audio. Veo generates sound as part of the same pass that produces the picture, so the clip arrives with its own atmosphere rather than as a silent file waiting for a sound library. None of the other video models here do that. If you have ever spent longer finding a matching sound effect than you spent making the shot, that is the problem this solves.

It is worth being precise about what that sound is and is not. It is the audio belonging to the shot: the room, the weather, the object, the movement. It is not your narration. The voiceover for a video is generated separately in MarsClip, in any of 32 languages, and captions are timed word by word against that voiceover. Veo gives a scene presence; the voiceover gives the video its argument. In practice they sit on top of each other.

The look is photoreal by default, which is a real preference rather than a neutral one. Veo is at its most convincing when it is pretending to be a camera: natural light, plausible physics, a lens that behaves like glass. Push it towards heavy stylisation and you are working against its instincts, and one of the other models will usually get there faster.

The twelve-second ceiling is a per-shot limit, not a limit on your video. Finished videos here are assembled from many generated scenes against a script, so the length of the piece is set by what you wrote. The ceiling only matters when a single moment genuinely must not be cut. That is the case where you would reach for Seedance 2.0 instead.

Veo 3.1 vs the other models here

The five models available in MarsClip are not five attempts at the same thing. Each is better at something specific, and choosing well is mostly a matter of naming which constraint your shot actually has: sound, length, character continuity, or speed.

Veo 3.1 vs Seedance 2.0 is the length trade. Veo runs to twelve seconds; Seedance runs to twenty-four, and is the longest single take available here. Some moments have to hold without a cut: a slow push across a landscape, a process that has to be seen through, a beat where cutting away would break the point. For those, Seedance is the answer and you accept a silent clip. If twelve seconds covers the shot, Veo gives you those seconds with audio attached, which is usually the better deal because most shots in a narration-led video are shorter than twelve seconds anyway.

Veo 3.1 vs Kling 3 is the continuity trade. Kling holds a character across a cut, which is the hard part of any video where the same person, creature or product appears in shot one and again in shot four. Veo is not built around that promise. If your video has a recurring subject and viewers will notice them changing, generate those shots on Kling and use Veo for the scenes that stand alone: the establishing wide, the atmospheric insert, the cutaway with its own sound.

Veo 3.1 vs Grok Video is the speed trade. Grok, from xAI, is the fast one, doing both image-to-video and text-to-video, and it is the sensible choice when you are trying ten ideas rather than finishing one. Veo asks for more patience and returns more polish. A reasonable working pattern is to rough a sequence out on Grok, find which shots the video actually needs, then regenerate the two or three that carry weight on Veo.

Veo 3.1 vs Nano Banana 2 is not really a comparison, because Nano Banana 2 is an image model rather than a video one. It generates and edits stills up to 4K, and the still you animate very often starts there. Set the frame you want on Nano Banana 2. Composition, colour, the exact product or face. Then bring it into a video model as a starting image. That route trades some of Veo's freedom for control over what is in shot, and it is the right route whenever the subject is fixed rather than invented.

The short version: reach for Veo 3.1 when the shot wants to sound like somewhere. Reach for Seedance 2.0 when it must not be cut, Kling 3 when someone has to stay themselves, Grok Video when you need answers quickly, and Nano Banana 2 when the picture matters more than the motion. Because the model is chosen per scene, a single video can and often should use several.

Getting better results from Veo 3.1

  • Write the sound into the prompt

    This is the one habit that changes results most, and the one most people skip. Do not only describe what is in frame. Describe what you would hear standing there. "Wet tyres on tarmac, a bus passing off-screen" produces a different clip from the same shot described silently, and it is the entire reason to be on this model.

  • Describe the shot as a photographer would

    What is in frame, where the camera sits, what moves and in which direction. "Low angle, a hand setting a cup down on a steel counter, steam rising" can be rendered. "Cinematic" and "epic" describe how you hope to feel afterwards, and nothing can be done with them.

  • Let it be a camera

    Veo is most convincing when the scene is one a camera could plausibly have captured: real light, real materials, physics that behave. Ask for something photoreal and it will usually oblige; ask for heavy stylisation and you are fighting the model rather than using it.

  • One event per shot

    Twelve seconds holds a single action well and a sequence of three badly. Split a complicated beat into two prompts and cut between them rather than asking one clip to do the work of an edit.

  • Mind the mix

    A Veo clip arrives with its own audio, and a video assembled from several of them can end up with a busier soundtrack than you intended once narration sits on top. Ask for the sound the scene needs, not everything that could be in it.

  • Mix models within one video

    The model is picked per scene. Use Veo for the shots that want atmosphere, Kling for the ones with your recurring character, Seedance for the one long unbroken take. A video built this way is better than one built by picking a favourite and forcing every shot through it.

  • Check the vertical crop

    A wide photoreal shot composed for a laptop can lose its subject when cropped for TikTok or Reels. Frame with the vertical cut in mind, and check captions on a phone-sized frame before anything is scheduled.

  • Fix the script first

    A convincing generated shot cannot rescue a thin script, on this model or any other. Get the words right, break them into scenes, and generate against them. The pictures are there to serve a line, not the other way round.

The other models available

  • Kling 3Cinematic motion that holds a character across a cut. The one for a recurring subject.
  • Seedance 2.0ByteDance, up to 24 seconds. The longest single take here, for shots that must not be cut.
  • Grok VideoxAI. Fast image-to-video and text-to-video. The one for trying ideas quickly.
  • Nano Banana 2Google's image model. Generation and editing to 4K. Where the still you animate usually starts.

Questions

Does Veo 3.1 really generate sound?
Yes, and it is the only model available here that does. Picture and audio come out of the same generation, so the clip arrives with the atmosphere of the scene rather than silent. Describing that audio in your prompt is what makes the difference.
How long can a Veo 3.1 clip be?
Up to twelve seconds per generation. That is a limit on the individual shot, not on your video. Finished videos are assembled from many scenes against a script, so the length of the piece is set by what you wrote. When one moment must hold longer than twelve seconds without a cut, Seedance 2.0 runs to twenty-four.
Is the generated sound the same as the voiceover?
No. Veo generates the sound belonging to the shot. Weather, room tone, movement, objects. Narration is separate: voiceover is generated in any of 32 languages and captions are timed word by word against it. The two sit together in the finished video.
Can I start a Veo shot from an image?
Yes. Describing a shot in words and starting from a still are both supported routes to a clip. Starting from an image is the better one when the subject is fixed and has to keep its real shape and colour. Stills can be generated or edited on Nano Banana 2 up to 4K first.
Should I use Veo 3.1 for every shot in a video?
Usually not. The model is chosen per scene, so most good videos mix them: Veo where atmosphere matters, Kling 3 where a character recurs, Seedance 2.0 for a long unbroken take, Grok Video when speed matters more than polish.
What can Veo 3.1 not do well?
It is built to look like a camera, so heavy stylisation is not where it is strongest. It makes no promise about holding a character consistent across separate generations, and it stops at twelve seconds. Each of those has a better answer among the other models here.
Where does the finished video get rendered?
With Remotion on AWS Lambda, not on your machine, so a long video does not tie up your laptop while it assembles.
Can it publish the finished video for me?
Yes. Connect YouTube, TikTok and Instagram and set a schedule; series and automation can keep a channel posting without you opening the app.
Can I use Veo 3.1 output commercially?
What you make is yours. Check the terms for the specifics before you build a campaign or a client deliverable on it.

Veo 3.1: Google's Video Model with Sound

Veo 3.1 is Google's video generation model. It produces photoreal shots of up to twelve seconds, and unlike every other model here it generates sound along with the picture. Run it inside MarsClip and the shot lands in a project that already handles script, voiceover, captions and publishing.

Generate with Veo 3.1