Gemini API Pricing: How Much Is 1M Input and Output Tokens?

From Qqpipi.com
Jump to navigationJump to search

The Gemini API has become a popular choice for developers and enterprises seeking large language model (LLM) capabilities with flexible pricing options. As of August 2026, the pricing landscape and plan ladder have seen some intriguing changes including recent renames and notable price cuts. This blog post provides a clear breakdown of Gemini API pricing, focusing on the cost of 1 million input and output tokens under the latest plans. We'll also analyze key themes such as usage limits, features like Deep Research and Flow credits, storage considerations, and the split of the Ultra tier into 5x and 20x performance levels.

Understanding Gemini API Pricing Structure

Before diving into specific costs, let's clarify the basics of the Gemini API pricing. The backbone of the pricing depends on token consumption, which is billed separately for input and output tokens. Additionally, cache hits on previously processed inputs can reduce input token costs significantly — a crucial factor for optimizing expenses.

Input vs Output Cost

API users often confuse input and output token costs, but they matter distinctly. Input tokens are the tokens in the prompt you send to the model, while output tokens are tokens generated by the model in response. Typically, output tokens are more expensive or priced similarly because they represent the computation and generation effort.

The Gemini API charges per million tokens, with specific rates varying by plan tier. Some tiers offer discounted input pricing for cached tokens, meaning if an input prompt was processed before, you pay less or nothing for repeating it.

August 2026 Gemini Plan Ladder and Pricing

As of August 2026, Gemini's pricing structure is segmented across Free, Standard, Plus, and Ultra tiers. Most notable is how the Ultra tier has splits into 5x and 20x variants, corresponding to computational performance and cost scales. Here is the current pricing ladder with relevant billing details:

Plan Monthly Price Input Tokens Price (per 1M) Output Tokens Price (per 1M) Comments Free $0 $0 $0 Limited usage, ideal for exploration Standard $49 $0.30 $0.60 Entry-level access with basic features Plus $149 $0.20 $0.45 Includes Deep Research feature Ultra 5x $499 $0.08 $0.20 5x performance, supports Flow credits and storage Ultra 20x $1,299 $0.05 $0.12 Top-tier speed, maximum throughput

Takeaway:

  • Free tier covers minimal token usage with no charges.
  • Input is roughly half the output cost in paid plans.
  • Ultra tier’s split caters to different throughput demands with significant price differentiation.

Recent Renames and Price Cuts

Gemini’s pricing has evolved significantly since its debut. Most notably, the Ultra tier's split — unveiling Ultra 5x and Ultra 20x — is a fresh development aimed at addressing diverse customer needs. Previously, the Ultra plan was a single, high-cost EU DMA decision 2026-07-27 tier, which caused some confusion and frustration among power users.

Price cuts have also been introduced to the Standard and Plus tiers, reducing token rates by nearly 20%, a move reflecting Gemini’s push to attract more volume-based usage. Renames in the Deep Research feature and Flow credits help clarify what resources are allocated per plan and better differentiate plan features from pure consumption limits.

Usage Limits vs Features: Deep Research, Flow Credits, and Storage

It’s important not to confuse pure usage limits (token counts) with feature availability:

  • Deep Research: Available starting in the Plus tier, this feature enables more sophisticated query handling and in-depth content searches that generate additional tokens but provide higher model accuracy.
  • Flow Credits: Exclusive to Ultra plans, Flow credits allow burst handling of high-complexity tasks without extra billing at times.
  • Storage: Persisted session context storage is bundled with Ultra tiers, allowing you to cache tokens and context across API calls for cost efficient input reuse and quicker responses.

Standard and Plus tiers offer minimal or no storage options, pushing power users seeking large scale integrations toward Ultra plans.

Cached Input Pricing

One insidious gotcha with API usage is repeated inputs. Gemini recently introduced cached input pricing, meaning if you submit an identical prompt that has been seen within your retention window, input tokens are charged at a discounted rate or zero cost based on the plan. This can significantly reduce cost for applications where queries repeat but outputs may vary slightly or be cached.

Calculating the Cost of 1M Input and Output Tokens

Let's break down a few examples to clarify exactly how much you pay for 1 million input and output tokens under key plans:

  1. Free Plan: 1M input + 1M output tokens cost $0 total. However, the free tier has strict token limits, so this is mostly theoretical.
  2. Standard Plan:

Token TypeCost per 1M tokensTotal Cost per 1M Input$0.30$0.30 × 1M = $0.30 Output$0.60$0.60 × 1M = $0.60 Total $0.90

  1. Plus Plan:

Token TypeCost per 1M tokensTotal Cost Input$0.20$0.20 Output$0.45$0.45 Total $0.65

  1. Ultra 5x:

Token TypeCost per 1M tokensTotal Cost Input$0.08$0.08 Output$0.20$0.20 Total $0.28

  1. Ultra 20x:

Token TypeCost per 1M tokensTotal Cost Input$0.05$0.05 Output$0.12$0.12 Total $0.17

Summary:

If you consume 1M input and 1M output tokens, expect to pay as little as $0.17 on Ultra 20x or up to $0.90 on the Standard plan per million tokens processed. Most real-world applications fall somewhere in between depending on plan benefits and caching.

Putting It All Together: The Pricing Bottom Line

  • Choose plans based on throughput and feature need: Standard and Plus are cost-effective for moderate use, while Ultra is tailored for mission-critical applications needing high throughput, storage, and Flow credits.
  • Remember cached input pricing: Aggressively reuse prompts where possible to benefit from discounted input charges.
  • Watch your token mix: Output tokens generally cost more, so optimizing prompt length versus desired output size can cut costs.
  • Stay current: Gemini’s aggressive price cuts and tier renames in 2026 mean outdated pricing info can mislead procurement decisions.

Final Thoughts

Understanding how much 1 million input and output tokens cost on Gemini API in 2026 is critical to managing budgets and architecting cost-efficient applications. The split of Ultra into 5x and 20x tiers allows a fine-tuned balance between speed and cost. Recent price cuts and clearer plan features like Deep Research and Flow credits make Gemini’s ecosystem more accessible and flexible.

If you’re evaluating or optimizing your Gemini usage, keep these nuanced pricing details in mind and leverage caching strategies. Avoid quoting stale high-tier prices from old Ultra plans ($249.99 Ultra) — they no longer reflect today’s cost efficiencies.

For continuous updates on Gemini and other SaaS pricing, keep an eye on pricing changelogs and always sanity-check storage bundles versus actual spending.