Skip to content
LaFoto

Google · model directory

What Is Nano Banana? Google’s image editing model explained

Nano Banana is the nickname for Google’s Gemini image model, built for conversational editing that keeps a subject consistent across changes.

Nano Banana is the nickname for Google’s Gemini image generation and editing model. Its distinguishing feature is conversational editing: you supply an image, describe a change in plain language, and it applies that change while keeping the rest of the picture — particularly a person’s face — recognisably the same. The name was a community label that Google subsequently adopted.

Nano Banana at a glance

Maker
Google
Family
Gemini
Access
Gemini app and API
Best known for
Character-consistent editing
Weights
Closed
Interface
Conversational

Where the name came from

It was not a product name. An unlabelled model appeared in public head-to-head evaluation arenas, where people compare anonymous outputs without knowing which system produced them, under the placeholder "nano banana". It won a great deal, people noticed, and the guessing about who was behind it ran for weeks before Google confirmed it.

Google then did something unusual and kept the nickname, which is why a serious piece of infrastructure is publicly known by a joke handle. The formal identity is a Gemini image model; the searched-for name is the banana.

That history is worth knowing because it explains why the term behaves so oddly in search. Interest arrived before any marketing did, so the questions people ask are the ones you ask about a rumour — what is it, who made it, is it real — rather than the ones you ask about a product.

What it actually does differently

Most image models are built around generating a picture from a description. Nano Banana is built around changing a picture you already have, and the hard part of that is everything you did not ask it to change.

Ask a general-purpose model to put the person in this photograph in a red coat and you will usually get a person in a red coat who is not quite the same person — a slightly different jaw, different eye spacing, a face that reads as a sibling. Identity survives one edit and degrades over several. The thing Nano Banana is built to hold onto is that identity, across a run of successive instructions.

The interaction model follows from that. You are not writing one elaborate prompt and accepting what returns; you are having a conversation in which each turn adjusts the last result. Move the light. Now make it evening. Now put him outside. That loop is the product.

What it is good for, in practice

Anything where the same subject has to appear more than once. Product photography in several settings, a character across a sequence of panels, a set of variations on one portrait, a mock-up that has to show one room in four different colourways.

It is also unusually forgiving for people who do not want to learn prompt syntax. Because the instruction is an edit rather than a specification, plain sentences work, and the failure mode of a vague request is a small change rather than an unrecognisable picture.

Where it is a poorer fit is the blank-page case. When there is no source image and no subject to preserve, the consistency machinery has nothing to do, and models built for text-to-image composition tend to give you more control over framing, optics and style from a standing start.

How it sits against the alternatives

Against Midjourney, the divide is aesthetic authorship. Midjourney applies a strong house style and is generally used to make something striking; Nano Banana is trying to leave the picture alone except where you asked. One is a stylist, the other is a retoucher.

Against the open models — Stable Diffusion and FLUX — the divide is where the control lives. Open weights let people bolt on ControlNet, LoRAs and inpainting pipelines to get precise control at the cost of assembling the pipeline. Nano Banana puts a narrower slice of that control behind a sentence, which is faster to reach and harder to push past.

Against gpt-image-1, the two are closer than either is to anything else: both are closed, instruction-following, conversationally driven models attached to a large assistant. The practical difference most people report is which one holds a face better over several turns, and that is worth testing on your own material rather than taking from a leaderboard.

What this has to do with LaFoto

Nothing commercially — we are not affiliated with Google and do not resell its models. This page exists because people ask what the thing is and the honest answer is short, and a directory that only described the models we happen to like would not be a directory.

It is worth reading alongside the editing pages here, because the vocabulary transfers. What Nano Banana does conversationally, the open ecosystem does with named techniques: inpainting for a change inside a mask, outpainting for a change beyond the frame, denoise strength for how far a result is allowed to drift from the source. Knowing those words makes it much clearer what any editor is and is not doing.

Editing versus generating

The problem it was built to solve

Ask any general-purpose image model to change one thing about a photograph of a person and you usually get back a person who is almost right. The jaw has shifted, the eyes sit differently, the face reads as a relative rather than the same individual. One edit survives it; four in a row do not.

Holding identity constant while everything you asked for changes is a specific engineering problem, and it is the one this model is pointed at. That focus is why it is described as an editor rather than a generator even though it can do both.

The same portrait photograph printed four times with different lighting on each print

Three things it is genuinely used for

Where consistency across images is the requirement

The same object in six settings

A product photographed once, then placed on a kitchen counter, a café table, an outdoor bench and a studio backdrop, staying identifiably the same object throughout.

A generator asked for the same product six times returns six similar products. That difference is the entire commercial argument for an editing model.

  • Shoot once, place many times.
  • The object must survive the change of scene.
  • Check the label and the logo on every output.
An AI-edited product scene with a replaced background

Nano Banana questions

What people ask after the first session

The things that come up once the novelty has worn off.

Why does the face still drift eventually?

Each edit is conditioned on the previous result rather than the original, so small errors compound. Re-supplying the source image every few turns resets the reference and stops the slow slide.

Can I get the same result twice?

Not reliably. Conversational editing does not expose a seed the way an open pipeline does, which is the general trade for the interface being this simple.

Is it better than inpainting?

It is easier. Inpainting with an explicit mask gives you exact control over which pixels may change; a sentence gives the model that judgement. Precision work still wants the mask.

Can it work from more than one reference image?

Combining references is a common request and support varies by surface and version. Test it on your own material rather than relying on any third-party description.

What is it bad at?

Starting from nothing. With no source image the consistency machinery has nothing to preserve, and models built for text-to-image composition give more control over framing and optics.

Does it work on non-human subjects?

Yes, and often more reliably — a product or a building has fewer features a viewer scrutinises than a face does.

An AI-generated studio portrait with soft, wrapped lighting

Keep exploring

Elsewhere on LaFoto

Editing is one job. These are the others.

Nano Banana — questions people ask

What is Nano Banana?
It is the nickname, now used officially, for Google’s Gemini image generation and editing model. It is built for conversational editing that keeps a subject consistent across successive changes.
Who made Nano Banana?
Google. It is part of the Gemini family.
Why is it called Nano Banana?
It appeared anonymously in public model-comparison arenas under that placeholder name, performed strikingly well, and the label stuck. Google adopted it rather than fighting it.
Is Nano Banana open source?
No. The weights are closed and access is through Google’s app and API.
What is Nano Banana best at?
Editing an existing image while keeping the subject recognisable — the same face across several successive changes, which most image models lose after one or two.
Is Nano Banana better than Midjourney?
They do different jobs. Midjourney imposes a strong aesthetic and is used to create striking images from nothing; Nano Banana tries to change only what you asked and leave the rest of a real photograph alone.
Can Nano Banana generate images from text alone?
Yes, but that is not where its advantage lies. Its distinguishing capability is preserving a subject through edits, which requires a source image to preserve.
Does LaFoto use Nano Banana?
No, and we have no affiliation with Google. This page is reference material in our model directory.

Start creating today

Generate your first image with the best AI image generator.

Turn a sentence into a finished, photorealistic image in seconds — then refine every detail. No setup, no Discord, no GPU.

Join 4,200+ creators using LaFoto