Nano Banana 2.1: The Essentials

Last updated: October 8, 2026

Nano Banana 2.1 brings image generation, editing and video-guided stills into one Flash-tier model. Build a scene from natural language, combine up to 14 references, or turn a video clip into a new image. Panoramic and ultra-tall formats make it useful for game environments, banners and scrolling compositions.


Which Model Should I Use?

Nano Banana 2.1 is the model covered here. Start with text for a new scene, add reference images when appearance or existing artwork matters, and use a video when the desired image needs visual context from a clip. Reference images and a video are alternative inputs and cannot be submitted together.

This guide uses real outputs from a 40-example onboarding batch: 30 practical examples and 10 experimental tests. Experiments explore difficult constraints; they do not imply guaranteed exact reproduction.


How to Use the Model

Describe a complete visual scene

Specify the subject, setting, style, lighting, action and composition in natural language. Explain the intended deliverable, such as a flat event poster, studio product photograph or scrolling game background. For text, put the exact words in quotation marks and describe their placement and typography.

Moonlight Market poster

A finished illustrated poster with two specified text lines. The title and time were legible in this output, with no prompt instructions rendered as captions.

Moonlight Market poster

Open this asset in Scenario

Create a polished illustrated poster for a fictional night market called "MOONLIGHT MARKET". A tiny lantern-lit street winds between indigo buildings, with orange citrus stalls, a friendly cream-colored cat and a cobalt night sky. The only lettering is "MOONLIGHT MARKET" in large elegant cream serif capitals at the top and "FRIDAY 18:00" in smaller clear capitals at the bottom. Keep both lines fully legible with generous margins, balance the detailed street against uncluttered typography, and use warm amber light, subtle screen-print grain and a sophisticated teal, navy and apricot palette. This is a finished vertical event poster, not a photograph of a poster, with no extra captions or watermarks.

Settings: 2:3, 4K, HIGH thinking, search off.

Compose for extreme formats

Select a wide or tall aspect ratio and describe how the scene should use that space. For a game environment, specify the ground position, depth layers and distribution of landmarks. For a vertical environment, explain how the zones connect from top to bottom.

Cloud orchard side-scroller

A floating orchard stretches across the panoramic format. The returned file measures 11712 × 1408, approximately 8.32:1, despite the 8:1 setting. Check dimensions before integrating the image into a fixed-size layout.

Cloud orchard side-scroller

Open this asset in Scenario

A hand-painted 2D side-scroller background of a floating orchard above a sea of clouds, with peach trees, weathered bridges, tiny windmills and waterfalls falling into empty sky. Establish a continuous walkable ground strip along the lower quarter, keep foreground branches sparse, separate near islands from distant mountains through atmospheric color, and distribute interesting landmarks across the entire panorama instead of clustering them at the center. Use luminous morning light, soft cel shading and a restrained jade, coral and cream palette; include no characters, interface, lettering or borders.

Settings: 8:1, 4K, HIGH thinking, search off.

Vertical underwater expedition

A single underwater environment connects the sunlit reef to deep vents. The returned file measures 1408 × 11712, approximately 1:8.32, despite the 1:8 setting.

Vertical underwater expedition

Open this asset in Scenario

A continuous ultra-tall illustrated underwater expedition from a sunlit coral reef at the top to a dark hydrothermal vent at the bottom. A small yellow research submarine travels through the middle depth, schools of silver fish occupy the upper water, a translucent jellyfish drifts below, and pale tube worms surround the deepest vent. Connect the zones through a gradual blue-to-black lighting transition with believable scale, generous breathing room and readable silhouettes. Make one cohesive vertical environment for a scrolling game, not stacked separate panels; include no typography or interface.

Settings: 1:8, 4K, HIGH thinking, search off.

Edit an existing image

Supply a reference image and state what changes and what must remain. Name the details that matter, including label text, materials, geometry, camera position and margins. A visually close edit is not a promise of pixel-perfect registration.

Tea packaging studio

The original sage matcha tin provides the reference for the color edit below.

Tea packaging studio

Open this asset in Scenario

Create a premium studio photograph of a single freestanding rectangular matcha tea tin, sage green with a ivory paper label and a brass lid. The label reads exactly "MOSS & MIST", "CEREMONIAL MATCHA" and "30 g", with crisp centered serif branding and small clean secondary type. Place the tin on warm limestone beside two tea leaves, lit by a large softbox from camera left with a gentle grounded shadow, slight surface texture and an uncluttered beige background. Show a three-quarter view with the entire tin visible, natural proportions and legible lettering; no additional products or text.

Settings: 4:5, 4K, HIGH thinking, search off.

Matcha tin label-preserving recolor

A targeted recolor changes the tin to cobalt blue while retaining its label, brass lid, limestone surface and leaves in this result. Compare both outputs before using a product variation commercially.

Matcha tin label-preserving recolor

Open this asset in Scenario

Edit the supplied photograph of the MOSS & MIST matcha tin on limestone. Change only the painted metal body from sage green to deep matte cobalt blue. Preserve the brass lid, the ivory paper label, every letter of "MOSS & MIST", "CEREMONIAL MATCHA" and "30 g", the two tea leaves, the tabletop, lighting, camera angle, object scale and margins. Keep the original tin geometry and grounded shadow, and maintain natural photographic material response. The result is the same product photograph with a precise color variation, not a redesigned package or a different scene.

Settings: 4:5, 4K, HIGH thinking, search off, 1 reference image(s).

Combine multiple references

Assign each input a clear role and refer to it by its order in the uploaded list. Distinguish the character, product, artwork or environment supplied by each image. The following experimental scene uses all 14 reference slots and gives each reference a separate place in an exhibition.

Fourteen-reference exhibition

The exhibition preserves recognizable subjects from the supplied references. Small reproduced text and exact object scale still need close inspection. This is an experimental composition, not an exact-copy guarantee.

Fourteen-reference exhibition

Open this asset in Scenario

Build one coherent contemporary design exhibition using all fourteen supplied references, in the exact supplied order. Reference 1 supplies a framed MOONLIGHT MARKET poster; 2 supplies the sage MOSS & MIST tin on a plinth; 3 supplies the cream courier robot with copper faceplate, teal backpack and orange boots; 4 supplies a wide framed aurora train artwork; 5 supplies a BICYCLE PARTS print; 6 supplies the PETIT FOUR menu; 7 supplies a screen showing My Garden; 8 supplies an ammonite exhibit; 9 supplies a MIDNIGHT NOODLES banner; 10 supplies a nasturtium botanical print; 11 supplies the ivory toy astronaut; 12 supplies the SOLAR ENERGY print; 13 supplies the Portuguese coffee poster; 14 supplies the Japanese stationery poster. Show exactly fourteen separately identifiable exhibits across a spacious warmly lit museum room, preserve each reference identity and distinctive colors, give each object a clear separate position without blending identities, and unify everything through one believable eye-level camera, warm white walls and subtle floor reflections. Do not add new labels or invent additional exhibits.

Settings: 21:9, 4K, HIGH thinking, search off, 14 reference image(s).

Use a video as context for a still

Supply one video instead of reference images. Describe the subject and moment you want the new image to represent. The model generates a still guided by the clip; it is not documented as a frame-extraction tool that returns an unchanged source frame. Scenario accepts a video of approximately 15 MB or less and exposes a frame sampling rate.

Video-to-image chef thumbnail

A new cinematic still guided by a five-second robot-chef clip, sampled at 2 fps. The source clip is separate from the 40 generated images.

Video-to-image chef thumbnail

Open this asset in Scenario

Use the supplied video of a dark metal multi-armed robot chef cooking in a neon-lit street food stall as visual context. Create a sharp cinematic thumbnail showing the same round-headed robot with amber glowing eyes, working above a black wok with the two upper arms holding blue-flame cooking torches and the lower arms handling utensils. Preserve the stall layout, hanging pans, wet teal city background, warm orange burner and brushed metal materials from the clip. Select a coherent expressive cooking moment rather than merging multiple times into ghosted limbs. Frame the chef and wok clearly, with no lettering, extra robot, new costume or exaggerated fire.

Settings: 16:9, 4K, HIGH thinking, search off, input video sampled at 2 fps.

Test structured layouts

Exact-count transit diagram

This experimental diagram renders twelve station dots and all twelve specified names. Count the objects and check the route connections independently rather than treating a successful generation status as proof of adherence.

Exact-count transit diagram

Open this asset in Scenario

Create a polished minimalist transit diagram on ivory paper titled "TWELVE STOPS". Arrange exactly twelve navy circular station dots in a single continuous serpentine route with three horizontal rows of four stations each, connected in reading order by one coral line. Put these exact station labels beside their corresponding dots in order: "Aster", "Birch", "Cedar", "Dune", "Elm", "Fern", "Grove", "Harbor", "Iris", "Juniper", "Kite", "Lake". The line begins at Aster and ends at Lake, with no branches or intersecting route, and every label appears exactly once. Use precise graphic geometry, consistent typography, generous spacing and no extra legend, logos or numbers.

Settings: 4:3, 4K, HIGH thinking, search off.


Parameters

prompt

Required natural-language description, up to 250,000 characters in Scenario's schema. Detailed scene descriptions with exact quoted text are used throughout these examples. See the Moonlight Market poster for text placement and typography.

referenceImages

Optional array of up to 14 images. Use clear inputs and explain the role of each. The exhibition example uses 14 references, while the matcha recolor uses one. Cannot be combined with video.

video

Optional input video, approximately 15 MB maximum. Use this instead of referenceImages. The robot-chef example demonstrates video-guided still generation.

videoFps

Sampling rate for an input video: 0.1 to 24 fps in increments of 0.1; default 1 fps. The chef example uses 2 fps. Other rates were not compared in this batch, so no quality advantage is claimed.

aspectRatio

Available presets: 8:1, 4:1, 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16, 1:4, 1:8 and auto. The batch covers every preset. The floating orchard and underwater examples show why actual output dimensions should be checked, especially for extreme formats.

resolution

Choose 1K, 2K or 4K. The botanical plate used 1K, the toy astronaut used 2K, and the poster and panoramic examples used 4K. These are resolution presets rather than a custom exact-width and exact-height interface.

useGoogleSearch

Optional Google Search grounding, off by default. It was enabled for the Fushimi Inari travel illustration and the James Webb telescope graphic in the batch. Search guidance does not remove the need to verify technical labels, geometry or factual accuracy.

thinkingLevel

Choose MINIMAL, MEDIUM or HIGH; default HIGH. The botanical plate used MINIMAL, the toy astronaut used MEDIUM and most examples used HIGH. These are different scenes, not a controlled speed or quality comparison.

numOutputs

Generate 1 to 4 images per request; default 1. Each showcase request in this batch used one output so prompts and assets remain individually traceable. Multi-output values were not tested.


Use Cases

Game development: wide environment concepts, character illustrations, inventory concepts and vertical levels. Marketing and commerce: product variations, event posters, packaging concepts and campaign banners. Film and content: cinematic stills and video-guided thumbnails. Education: diagrams and illustrated explanations that are checked by a subject expert.


Tips

Describe the composition in terms of the deliverable: a connected lower-third path for a side-scroller, an unobstructed product label, or three rows of four slots for a twelve-slot interface. Assign separate roles to references. Keep required text explicit and concise. Inspect the exported file dimensions rather than relying only on the preset name.

For precise grids and counts, name both the row count and column count. The first inventory output had twelve slots but arranged them as two rows of six rather than three rows of four, illustrating why the total count alone is insufficient.


Known Limitations

Aspect ratios are approximate in this batch. The 8:1 and 1:8 4K presets returned 11712 × 1408 and 1408 × 11712. Other presets also returned small deviations, so exact layouts may need a separate crop or resize step.

Text and decorative content can drift. The garden interface added a small TODAY label, and the botanical illustration added a Nasturtium caption despite instructions limiting the text. Complex layouts can satisfy an object count while missing the requested arrangement.

A multi-environment panorama can become separate panels rather than a continuous landscape. The experimental biome test was revised with explicit transition and ground-path instructions. Reference-based edits still require comparison of identity, scale and placement.

One benign bicycle infographic request was rejected with the provider's generic safety message; a simpler revised diagram succeeded. A rejection message alone does not establish which part of the request caused it.

Google Search grounding guides generation but does not certify scientific accuracy. Image references and video remain mutually exclusive. No direct provider-to-provider performance comparison was conducted.


Further Reading

Google DeepMind prompting guide, Gemini image generation documentation, and Google's Nano Banana overview. Scenario's live schema determines the controls available on this model.

Try Nano Banana 2.1 on Scenario.