New: the Nano Banana 2 Lite model is live on Nano Banana.

See the model
Back to all articles

img2img vs txt2img: a scenario-by-scenario field guide

Published 2026-07-19 · 6 min read

Text-only prompts invent; reference-anchored prompts preserve. When each mode wins, how reference images change what you should write, and how to mix both in one workflow.

The one-line difference

txt2img starts from words only — the model invents a subject to match them. img2img (reference-based generation) starts from your image plus words — the model preserves the subject and applies the change you describe. On Nano Banana both run in the same workbench: a batch carries 0–2 reference images, and every prompt in the batch shares them.

When txt2img wins

Use text-only when the subject itself is open: concept exploration, mood boards, background plates, ad hooks where the product will be composited later, or any case where surprise is the point. txt2img is also the right first step when you do not have a usable photo yet — generate candidate packshots, pick one, and that becomes your reference.

  • Concept and mood exploration.
  • Backgrounds and scene plates for later compositing.
  • Creative ad hooks where novelty beats fidelity.

When img2img wins

Use reference-anchored generation whenever the subject is fixed: real product photos that must keep their label and shape, brand assets, model shots that must stay on-model. The reference carries identity; your prompt carries only the delta — background, light, scene, season.

  • E-commerce listing and detail images from real packshots.
  • Seasonal reskins of the same product (same bottle, winter scene).
  • Background replacement and relighting of catalog photography.

How references change what you write

With no reference, completeness helps — subject, material, scene, light, mood. With a reference, completeness hurts: every word re-describing the product is a chance to drift from it. Write the change, not the object. And use the second reference slot only for information the first lacks — a texture swatch, a color card — never for a second angle of the same thing.

A mixed workflow that covers both

The strongest e-commerce loop uses both modes in sequence: txt2img to explore scene concepts cheaply at 1K, pick the winning scene, then img2img with the real product photo to re-render that scene on-model. Exploration stays inventive; production stays faithful.

  • Step 1 — txt2img batch: 20 scene concepts, pick 2.
  • Step 2 — img2img batch: real packshot + winning scene descriptions.
  • Step 3 — re-render winners at 2K/4K and ship.

The rule of thumb

If the image must stay true to something real, anchor it. If the image only has to be true to an idea, write it. Most production work is the former wearing the clothes of the latter.

Put it into practice.

Everything in this guide runs in the same workbench — open it and try the steps while they are fresh.

Open the generator