Nano Banana 2.1 arrives with 4K
The image model improves Chinese text rendering, adds 4K, and halves per-image cost: $0.050 at 2K and $0.076 at 4K.

Google released Nano Banana 2.1 (gemini-nano-banana-2.1) on October 6 — the newest generation of its image generation and conversational editing line, generally available immediately. Chinese AI blogger 歸藏 (Guizang, of the 歸藏的AI工具箱 newsletter) published one of the first hands-on takes: Chinese text rendering — clarity and typo rates — improves substantially, the Chinese typography “lacks the broken look GPT produces,” and the headline advantage over GPT-family rivals is native 4K output.
Key points
- Positioning: an upgraded Nano Banana 2 — image generation plus conversational editing at Flash-tier speed and cost
- Resolutions: 1K / 2K / 4K output (the 512px option remains exclusive to Gemini 3.1 Flash Image)
- Chinese text: clarity and typo rates improve markedly per first hands-on reports
- Pricing: $30 per million output image tokens — $0.0504 per 2K image, $0.0756 per 4K image, exactly half the previous generation
- Capabilities: editing with up to 14 reference images, Google Search and Image Search grounding, three Thinking levels, video-to-image, SynthID watermarking on every output
- API: served through the Interactions API at
generativelanguage.googleapis.com/v1beta/interactions
The arithmetic of the price cut
Google’s pricing page shows the mechanics plainly: per-image token consumption is unchanged (1K = 1,120 tokens, 2K = 1,680, 4K = 2,520), so the entire price drop comes from the output token rate — $30 against the previous generation’s $60. In creator terms: 4K images fall from $0.151 to $0.076, halving the cost of batch workflows in e-commerce and social media. The offset is on the input side, up 3x from $0.50 to $1.50 per million tokens — heavy multi-reference editing sessions need a fresh cost model.
Why Chinese text rendering matters
Chinese glyph rendering has been the durable weakness of image models — signage, posters and UI mock-ups full of stroke-noise that looks like characters but isn’t. Combining 14-reference editing with 4K output puts a complete poster-grade delivery workflow inside one model for the first time; Guizang’s hands-on calls Chinese usability the headline change in 2.1, which is a direct productivity shift for Chinese-market design tooling; combined with 4K output, generations move from “preview” to “deliverable.” The verification path is trivially reproducible: run identical prompts on 2.0 and 2.1 and compare stroke accuracy on signage text.
For buyers, the comparison that matters is per deliverable, not per token: at $0.076 per 4K image, a hundred-asset campaign costs $7.60 of output — the kind of number that moves image generation from procurement discussion to credit-card purchase. The open question is throughput under load and how the 3x input price compounds across long editing chains; Google’s docs do not publish latency figures, so the first week of production reports will settle it.
The competitive landscape
The image-generation price list is crowded: FLUX.3’s maximum-fidelity route, GPT-family native multimodality, Midjourney’s subscription, and Nano Banana’s token-metered conversational editing. 2.1’s differentiation is a three-part combination — iterative editing without rewriting prompts, Search grounding that pulls real references into generation, and Google Cloud metering, which for engineering teams often decides selection more than benchmarks do. The hands-on praise for work that lacks “the broken look” targets the classic failure of multi-turn editing: canvases that degrade as rounds accumulate.
Efficiency pricing, again
The Nano Banana line has always chased “Flash price, Pro quality,” and this adjustment — outputs halved, inputs tripled — reallocates cost toward high-frequency conversational editing, where multi-turn, multi-reference sessions make input the growing share. For developers, image-model competition is now contested at pennies per image; for everyone, SynthID watermarking on all outputs is the shared substrate of this model generation — provenance ships by default, whether or not your jurisdiction demands it yet.
One adoption note: the model is served through the new Interactions API rather than the classic generateContent surface, so existing pipelines need a small migration; the docs position interactions as the home for multi-turn editing state, which is where 2.1’s real product lives. Teams running batch generation can also use the Batch tier at half price — $0.038 per 4K image — for non-urgent workloads.