AI Image Generator
Describe a picture and get it back as an image, up to 4K. Here the still is rarely the end of the job. It is the frame you approve before you animate it, which is why generation and video sit in the same workspace.
How to generate an image
Write the description
Say what is in the picture, how it is lit, and how it is framed. "A ceramic kettle on a steel worktop, morning light from a window on the left, shallow depth of field" is a brief. "Beautiful kitchen photo" is a wish.
Set the shape before you generate
The aspect ratio sits beside the prompt, and the options follow the model you have chosen. Decide it at the start, not after: a 16:9 still cropped to a vertical post loses the sides of the composition, and the subject you carefully placed off-centre ends up half out of frame.
Generate a handful, then choose
The same prompt returns a different picture each time. Reading several results next to each other tells you far more about what the prompt is actually asking for than staring at one and hoping.
Refine, then carry it forward
Adjust the wording where the result missed, or take the image into the editor to change one part of it. When the frame is right, it goes to your library. From there it can go into image-to-video, or into an edit with voiceover and captions.
What people make with it
First frames for video
The most common use here. An image-to-video model inherits everything in the still it is given, so getting the picture right is the cheapest possible place to fix a shot. Images come back faster than video, which makes this the sensible place to be indecisive.
Faceless channel visuals
History, science, finance and story channels are built almost entirely from images. Generating them means you are not trawling stock libraries for a picture that half fits, and the visual style stays consistent from one video to the next.
Thumbnails and covers
A thumbnail has one job: be legible at the size of a postage stamp. Generating a few candidates at a high resolution lets you test the composition small before you commit the video to it.
Product and concept visuals
Mock up a packaging idea, a scene a product might sit in, or a variation on a key visual. Useful when you want to see the concept before anyone books a photographer.
Storyboards and shot planning
Generate the sequence as stills first: opening frame, mid shot, close-up, closing frame. Approving the boards before you animate anything is a much shorter feedback loop than approving finished clips.
Ad creative variations
One idea, several visual treatments. Because the stills come back quickly, you can put five directions in front of someone instead of defending your favourite one.
What is an AI image generator?
An AI image generator is a model that turns a written description into a picture. You type what you want to see: the subject, the setting, the lighting, the framing. The model produces an image that matches it, rather than retrieving one that already exists.
The models are probabilistic, which is the single most useful thing to understand about them. There is no lookup table and no correct answer being fetched: the same prompt produces a different picture every time you run it. That is why the working method is to generate several and choose, not to write one perfect prompt and expect one perfect result.
Resolution matters more than it first appears. An image at 4K survives being cropped, being pushed in on, and being used as the first frame of a video clip. A small, soft image looks acceptable on its own and falls apart the moment you ask anything of it.
In this product, generation is deliberately placed next to video. You are usually making the still because something is going to happen to it afterwards. It will be animated, edited into a sequence, or given a voiceover and captions. The tools sit together for that reason.
Generating an image vs editing one vs going straight to video
Three routes lead to a finished frame, and picking the wrong one is the usual reason a session goes slowly. Generation invents the picture from a description. Editing takes an image that already exists and changes part of it. Text-to-video skips the still entirely and asks a video model to invent the subject and the movement in one go. They are not ranked; they answer different questions.
Generate when nothing exists yet and the subject is not fixed. If you need "a lighthouse in fog at dawn" and no particular lighthouse is required, describing it is faster than sourcing it. Generation is also the right call when you want options: five interpretations of the same brief costs you one more click, where five photo shoots costs you a month.
Edit instead when the picture is nearly right. If you have a photograph of your actual product, generating a lookalike is the wrong move. The model will produce something in the same category, not the same object. Change the background, remove the distraction, adjust the light, and you keep the thing that mattered. The rule of thumb is simple: if a specific real object has to appear, start from a photograph of it.
Go straight to text-to-video only when you are exploring. It is the fastest way to see whether an idea has any life in it, because you are not obliged to produce a picture first. The trade is control: the video model decides what your subject looks like, and if it decides differently on the next generation, your two clips do not belong to the same film.
That last point is why the still-first workflow wins for anything longer than a single clip. Each video generation starts fresh, so consistency across shots is hard to hold with text prompts alone. Generate the frames as images, approve them, then animate each one. The clips come back looking like they were shot in the same place on the same day. Most people making a series arrive at this method after a few videos, usually the hard way.
A fourth option deserves naming, because it is genuinely the right answer sometimes: use a photograph you already have. AI generation is not a moral improvement on a camera. If the shot exists, or is easy to take, take it. Then use these tools for the things a camera cannot do, like animating it or producing eight variations of it before lunch.
Getting better results
Describe the picture, not the vibe
"Stunning", "epic" and "high quality" give the model nothing to place in the frame. Nouns and spatial relationships do: what is in the foreground, what is behind it, where the light comes from, what the camera is close to. Write it the way you would brief a photographer who cannot see what you are imagining.
Name the light
Lighting decides more of the final look than almost anything else you can specify. "Late afternoon sun through a window, long shadows" and "flat overcast daylight" produce two entirely different pictures of the same subject. If you say nothing about light, the model chooses for you, and it tends to choose the average.
Say where the camera is
Close-up, wide shot, eye level, from above, through a doorway. Framing is a decision, and leaving it unstated is how you get the same mid-distance centred composition every time.
Change one thing per attempt
When a result is close, resist rewriting the whole prompt. Alter one thing: the lighting, or the framing, or the background. Then generate again. Changing four things at once means the next result tells you nothing about which change helped.
Compose for what happens next
If the still is going to be animated with a push-in, do not let the subject fill the frame already. The movement needs somewhere to go. If it is going to carry a caption, leave a quiet area for the text. The frame you generate should suit the finished piece, not just look good on its own.
Generate high, downscale later
Going up in resolution after the fact is a reconstruction; starting high is just detail you already have. Since 4K is available, use it for anything that will be cropped, animated or reused. You can always make an image smaller, and the result is always clean.
Keep a house style written down
If a series should look consistent, write the style clause once: palette, lens feel, lighting, level of realism. Paste that same block into every prompt and vary only the subject. Consistency comes from repetition, not from asking for it.
Treat text in images with suspicion
Words rendered inside a generated picture are unreliable, and a misspelling on a thumbnail is worse than no words at all. Where the copy matters, add it in the edit rather than asking the model to draw it.
Models available
- Nano Banana 2 — Google's best stills. Generation and editing, to 4K.
- Kling 3 — Cinematic motion that holds a character across a cut. For animating the still you generated.
- Seedance 2.0 — ByteDance, up to 24 seconds. The long shots.
- Grok Video — xAI. Fast image-to-video, and text-to-video.
Questions
- What resolution can I generate at?
- Nano Banana 2 handles generation and editing up to 4K. Working at the higher end is worth it whenever the image will be cropped, animated, or reused at a different aspect ratio, because you are keeping detail rather than inventing it later.
- Can I choose the aspect ratio?
- Yes, and it is worth doing before you generate rather than cropping afterwards. The control sits next to the prompt and the available shapes follow the model you have selected. Composition is a decision about the edges of the frame, so changing the edges later undoes part of what you approved.
- Why does the same prompt give me a different image each time?
- Because the model is probabilistic. It samples a result rather than looking one up. This is a feature more often than a nuisance: generate several, pick the strongest, and use the differences between them to work out what your prompt is really asking for.
- Can I change part of an image instead of starting again?
- Yes, that is what the image editor is for. Regenerating from scratch when you only wanted to fix the background throws away everything that was already working.
- Can I turn the image into a video?
- Yes. That is the usual next step here. The still becomes the first frame of a clip, so the subject, colours and composition carry through into the motion instead of being reinvented.
- How do I keep a series of images looking consistent?
- Fix the style and vary the subject. Keep one written block describing palette, lighting and level of realism, reuse it in every prompt, and change only the part that describes what is happening. Approving the stills as a set before you animate anything also catches drift early.
- What can I do with the image after it is generated?
- It lands in your library. From there it can be edited, upscaled, animated into video, or dropped into a sequence with a script, voiceover and captions. Finished videos can be published to YouTube, TikTok and Instagram on a schedule.
- Can I use the images commercially?
- What you make is yours. Check the terms for the specifics before you build a campaign on it.
AI Image Generator
Describe a picture and get it back as an image, up to 4K. Here the still is rarely the end of the job. It is the frame you approve before you animate it, which is why generation and video sit in the same workspace.
Start creating