Nano Banana 2: Google's Image Model, Generation and Editing to 4K

Nano Banana 2 is Google's image model. It generates stills and edits existing ones at up to 4K. It is the only image model in the set, which makes it where most good shots begin. The still you animate is easier to get right than the video you generate blind.

How to use Nano Banana 2

  1. Start from a script or a topic

    Bring a script or have one written from a topic, and the video is broken into scenes before anything is generated. Knowing which shots the video needs tells you which stills are worth making carefully.

  2. Generate the still, or bring one to edit

    Describe the frame and generate it, or supply an image you already have and change it. Editing is the half people underuse: adjusting a photo you own often gets you to the right frame faster than inventing one from nothing.

  3. Iterate on the picture, not the video

    This is the whole argument for working image-first. Fixing a composition, a colour or a stray detail is far cheaper on a still than it is on a clip. Get the frame right here, where changing your mind is inexpensive.

  4. Animate it

    Take the finished still into a video model as a starting frame. Veo 3.1 when the shot wants its own sound, Kling 3 when a character has to recur, Seedance 2.0 for a long unbroken take, Grok Video when you want the motion quickly.

  5. Narrate, caption and publish

    Voiceover is generated in any of 32 languages, captions are timed word by word against it, and the video is rendered with Remotion on AWS Lambda. Connect YouTube, TikTok and Instagram to publish on a schedule, or run a series so a channel keeps posting without you opening the app.

What Nano Banana 2 is good for

Setting a look before anything moves

Colour, light, composition, the grade the whole video will live in. Deciding those on one still and then holding to it across every shot is what makes a set of clips read as a single piece rather than a folder of generations.

Pinning down a character

A character described in words drifts between prompts; a character fixed as an image does not. Generating them once at 4K and using that file as the starting frame for each shot is the most reliable route to continuity.

Editing product photography

You already own the shot of the product; what you need is a different background, a cleaner surface or a seasonal treatment. Editing keeps the object's real shape, colour and branding, which invention will not.

Thumbnails and cover frames

A thumbnail is a still doing a job on its own, at a size where detail matters. Generating one at 4K and cropping down beats pulling a frame out of a video and hoping it holds up.

Storyboarding a video cheaply

Six stills tell you whether a sequence works for a fraction of the effort of six clips. If the video does not survive as a set of pictures, it will not survive as footage either.

Fixing one thing that is wrong

A frame that is nearly right is a common outcome and a frustrating one. Editing addresses the specific problem. The sign that says the wrong thing, the object in the corner. Fix that rather than rolling the dice on a whole new image.

What is Nano Banana 2?

Nano Banana 2 is Google's image model. It does two things: it generates a still from a written description, and it edits a still you supply. Both run at up to 4K. Inside MarsClip it is the only image model in a set otherwise made of video models, which is why it occupies a particular position in the workflow rather than competing with the others directly.

The editing half is the part worth dwelling on, because it is the one people skip. Generation invents a frame from nothing; editing takes a picture that exists and changes it. If you have product photography, a location shot or a character you have already settled, editing preserves what is real about it and changes only what you asked for. Invention has no such anchor, and it is why a generated product never looks quite like your product.

The 4K ceiling matters for two reasons. The first is obvious: a thumbnail or a cover frame is looked at closely, and detail survives cropping only if it was there to begin with. The second is less obvious but more useful. A still you intend to animate is going to be examined for the whole duration of the shot, so the flaws you can ignore at preview size become the thing the viewer is looking at for twelve seconds.

That leads to the reason this page exists in a set otherwise about video. Generated video is hard to steer, because you are describing motion and content at the same time and can only judge the result after the clip exists. Splitting the problem in two makes it tractable: settle what is in frame as a picture, where iterating is cheap and the result is immediately legible, then ask a video model to do one thing. Move it. The still you animate usually starts here, and shots made this way go wrong less often than shots described from scratch.

One clarification on scope. This is an image model; it does not generate video, and it does not generate sound. Its output is a frame. What that frame is for is up to you: a thumbnail, a storyboard panel, a character reference, or the opening frame of a shot. In a video project it is most often the last of those.

Nano Banana 2 vs the video models here

The comparison here is unusual, because Nano Banana 2 is not an alternative to the other four models. It is the step before them. The four video models each hold a different constraint: sound, character continuity, length, and speed. Nano Banana 2 holds the frame. Understanding which of the five you need starts with noticing that "what is in shot" and "how it moves" are separate questions, and only one of them is a video question.

Nano Banana 2 with Veo 3.1. Veo, from Google, is the only model here that generates audio along with the picture, and it runs to twelve seconds. It is at its strongest when it is pretending to be a camera. Giving it a photoreal still as a starting frame plays directly to that: the frame is settled, so the generation is spending its effort on motion and atmosphere rather than on inventing a scene. Describe the sound in the prompt and you have a shot that both looks and sounds like somewhere.

Nano Banana 2 with Kling 3. This is the pairing that matters most. Kling is built to hold a character across a cut, and the most effective way to help it is to stop relying on words. Generate your character once on Nano Banana 2 at 4K, keep that file, and start every Kling shot from it. Text descriptions drift between prompts no matter how carefully you write them; a reference image does not. If continuity is your problem, this combination is the practical answer to it.

Nano Banana 2 with Seedance 2.0. Seedance, from ByteDance, runs to twenty-four seconds. That is the longest single take available here. Length amplifies whatever is wrong with the opening frame: a subject that is nearly right is tolerable for two seconds and painful for twenty-four. Fixing that frame as a still before committing to a long generation is a straightforwardly better use of your time than generating twenty-four seconds twice.

Nano Banana 2 with Grok Video. Grok, from xAI, is the fast one and does image-to-video as well as text-to-video, which makes this the most efficient loop in the set. Iterate on the picture where iteration is cheap, then animate it where animation is quick. Work through a long shot list that way. Settle the frame, move the frame, next. It beats describing each shot in words and waiting to find out what you got.

It is worth saying plainly where this does not apply, because the image-first argument is easy to over-extend. A great many shots do not need a still first: an atmospheric insert, weather, a crowd, a cutaway whose exact contents nobody will examine. Fixing a frame for those is an extra step that buys nothing, and describing the shot straight to a video model is quicker. Pinning the frame also costs something real. A model asked to invent the whole scene has room to return something you would not have thought to specify, and occasionally that is better than what you had in mind. Veo 3.1 in particular is at its most convincing when it is composing the shot as well as moving it, so handing it a fixed frame trades away part of what it is good at.

The short version: use Nano Banana 2 when what is in the frame matters more than what it does. A real product, a specific place, a character who has to stay themselves. Skip it when the shot only has to be plausible. Then reach for Veo 3.1 for sound, Kling 3 for a recurring character, Seedance 2.0 for a take that must not be cut, and Grok Video when speed is the thing you need.

Getting better results from Nano Banana 2

  • Describe the frame like a photographer

    What is in shot, where the camera is, what the light is doing. "Overhead, a chipped enamel mug on a dark wooden table, morning light from the left" can be rendered. "Beautiful" and "cinematic" describe how you hope to feel about the result, and nothing can be done with them.

  • Edit rather than regenerate when it is nearly right

    A frame that is ninety per cent correct is a good starting point, not a failure. Changing the one thing that is wrong keeps everything you already liked; rolling a new generation throws it all away and rarely lands closer.

  • Keep your character reference and reuse the file

    Generate the character once, properly, and save that image. Starting every shot from the same file is far more effective for continuity than re-describing them, and it costs one step rather than a rewrite.

  • Compose for the crop you will actually publish

    A frame composed for a wide format loses its subject when cropped to vertical for TikTok or Reels. Decide the aspect ratio before you generate, and leave room where captions will sit.

  • Leave somewhere for the motion to go

    A still destined to be animated should not be perfectly centred and perfectly closed. Space in front of a subject, a path for the camera, something entering frame. Each gives the video model an obvious move to make.

  • Use the resolution when it is going to be looked at

    Thumbnails, cover frames and anything held on screen for several seconds are examined closely. Generate large and crop down rather than generating small and enlarging.

  • Storyboard before you generate footage

    Make the six stills first and look at them in order. If the sequence does not work as pictures it will not work as clips, and you will have found that out for a fraction of the effort.

  • Fix the look once and reuse it

    The same palette, the same light, the same framing habits across every shot are what make a run of clips read as one channel. Decide those on one still, write down what you asked for, and stop relitigating it each time.

The video models available

  • Veo 3.1Google. Photoreal generation with sound, up to 12 seconds. The only model here that generates audio.
  • Kling 3Cinematic motion that holds a character across a cut. Pair it with a still to keep a subject consistent.
  • Seedance 2.0ByteDance, up to 24 seconds. The longest single take here, for shots that must not be cut.
  • Grok VideoxAI. Fast image-to-video and text-to-video. The quickest route from a still to a moving shot.

Questions

Is Nano Banana 2 a video model?
No. It is Google's image model. It generates and edits stills at up to 4K. Its place in a video project is the frame you animate afterwards with one of the video models.
Can it edit an image I already have?
Yes, and that is often the better route. Editing keeps what is real about a photo: a product's actual shape, colour and branding, a location as it looks. It changes only what you asked for. Generating from nothing has no such anchor.
What resolution does it produce?
Up to 4K. That matters for thumbnails and cover frames, which are looked at closely, and for any still you intend to animate, since a viewer will be looking at that frame for the whole duration of the shot.
Why generate a still before generating a video?
Because it splits one hard problem into two easy ones. Settling what is in frame is cheap and immediately legible on a picture. Asking a video model to do a single thing, move it, goes wrong far less often than asking it to invent the scene and the motion at once.
How does this help with keeping a character consistent?
Words drift between prompts; a reference image does not. Generate the character once at 4K, keep the file, and start each shot from it. Paired with Kling 3, which is built to hold a character across a cut, this is the most reliable approach available here.
Which video model should I animate the still with?
Veo 3.1 when the shot wants its own sound, up to twelve seconds. Kling 3 when a character recurs across cuts. Seedance 2.0 when the take must run up to twenty-four seconds without cutting. Grok Video when you want the motion quickly. The model is chosen per scene, so one video can use several.
Can I use the images on their own?
Yes. A still is a finished thing: a thumbnail, a cover frame, a storyboard panel. Nothing obliges you to animate it.
Where does the finished video get rendered?
With Remotion on AWS Lambda rather than on your machine, so assembling a long video does not tie up your laptop.
Can I use Nano Banana 2 output commercially?
What you make is yours. Check the terms for the specifics before you build a campaign or a client deliverable on it.

Nano Banana 2: Google's Image Model, Generation and Editing to 4K

Nano Banana 2 is Google's image model. It generates stills and edits existing ones at up to 4K. It is the only image model in the set, which makes it where most good shots begin. The still you animate is easier to get right than the video you generate blind.

Generate with Nano Banana 2