Justin McKelvey

Justin McKelvey

Fractional CTO · 15 years, 50+ products shipped

• AI for Business • 5 min read •

Gemini API Pricing (October 2026): Every Model, the January 1 Price Doubling, and What a Real Workload Costs

Gemini API pricing: the short answer

As of October 4, 2026, Google's rate card prices Gemini 3.8 Flash, its newest Flash model, at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. On January 1, 2027 that becomes $1.50 and $7.50. Gemini 3.1 Pro Preview is $2 and $12. Batch requests are half price. There is a free tier, but on it Google uses your content to improve its products. Priced on one real workload, Flash is less than half of Claude Sonnet 5.5 or GPT-6.1 Sol today, and still cheaper after the increase.

The headline everyone else skips is the date. Google has already published next year's price, and it is double this year's. If you are choosing a model this quarter, you are choosing at a promotional rate.

Every Gemini text model, one table

Paid tier, Standard requests, per million tokens, read from Google's pricing page today (it says it was last updated October 1):

Model Input Output From Jan 1, 2027 Free tier?
Gemini 3.8 Flash (newest) $0.75 $3.75 $1.50 / $7.50 Yes
Gemini 3.7 Flash $0.75 $3.75 $1.50 / $7.50 Yes
Gemini 3.6 Flash $0.75 $3.75 $1.50 / $7.50 Yes
Gemini 3.5 Flash $1.50 $9.00 No date listed Yes
Gemini 3.5 Flash-Lite $0.30 $2.50 No date listed Yes
Gemini 3.1 Flash-Lite $0.25 $1.50 No date listed Yes
Gemini 3.1 Pro Preview $2.00 ($4.00 over 200K) $12.00 ($18.00 over 200K) No date listed No
Gemini 2.5 Pro $1.25 ($2.50 over 200K) $10.00 ($15.00 over 200K) No date listed Yes
Gemini 2.5 Flash $0.30 $2.50 No date listed Yes
Gemini 2.5 Flash-Lite $0.10 $0.40 No date listed Yes

Output prices include thinking tokens, which matters more than it looks: a model that "thinks" before answering bills that reasoning as output. Gemini 4 Argon, announced September 30, isn't on the rate card yet; its intro and list prices, and how they line up against Claude, are in Claude vs Gemini.

The January 1 doubling, and the odd result it creates

Every 3.6, 3.7 and 3.8 Flash rate on the page carries two prices: one "through December 31, 2026" and one "starting January 1, 2027." Input goes from $0.75 to $1.50, output from $3.75 to $7.50, cached input from $0.075 to $0.15, and Batch, Flex and Priority all double with it.

That produces something you don't usually see on a price list. Right now the newest Flash is cheaper than the older one: Gemini 3.5 Flash is $1.50 in and $9.00 out, so 3.8 Flash costs half as much on input and well under half on output. On January 1 the input prices match at $1.50, and 3.8 Flash is still cheaper on output ($7.50 against $9.00). If anything you run is still on 3.5 Flash, moving it now saves money twice: once today, and again next year.

Two practical points. Budget 2027 at the January price, not today's. And if you are testing Flash this quarter to decide whether to build on it, multiply your test bill by two before you compare it to anything.

What one workload costs across Google, Anthropic and OpenAI

Rate cards are hard to compare because nobody's workload is "one million tokens." So here is one shape priced everywhere: 50,000 requests a month (an internal assistant or a support-triage tool at a small company), each with about 3,000 tokens in and 600 tokens out. That's 150 million input and 30 million output tokens a month, Standard tier, no caching, each price from the vendor's own page today.

Model Monthly cost Note
Gemini 2.5 Flash-Lite $27 Cheapest on the board
GPT-6 Luna $30 OpenAI's smallest current model
Gemini 3.8 Flash, Batch $112.50 $225 from January
Gemini 3.5 Flash-Lite $120
Gemini 3.8 Flash $225 $450 from January 1
Claude Haiku 4.5 $300
Gemini 3.5 Flash $495 Older and pricier than 3.8 Flash
Claude Sonnet 5.5 $600
GPT-6.1 Sol $600
Gemini 3.1 Pro Preview $660 No free tier
Claude Opus 5.5 $1,200

The spread is 44 to 1 from the cheapest line to the most expensive, on the same work. The full Claude and OpenAI rate cards, including caching and long-context rules, are on the Anthropic API pricing and OpenAI API pricing pages. The bill is the easy part; which of these actually does the job well enough is the question to test before you build, and it is most of what an AI implementation engagement spends its first week on.

Is the Gemini API free?

Partly. Google's own table lists a free tier with "limited access to certain models" and free input and output tokens, and says Google AI Studio usage "is free of charge in all available regions." Most Flash and Flash-Lite models have a free tier. Gemini 3.1 Pro Preview doesn't.

The real cost of free is one row in that same table: "Used to improve our products." On the free tier the answer is Yes. On the paid tier it is No. That's the line that should decide it for a business. Testing with public data on the free tier is fine. Customer emails, contracts or anything you'd mind seeing in a training set belong on the paid tier, even if the bill is a few dollars. The free tier's limits also change; Google's page says rate limits are "subject to change," so don't build a product on them. The separate question of how much you can use the Gemini app itself, on the consumer plans, is in Gemini usage limits.

The extras that add to the bill

  • Batch and Flex. Half the Standard price for work that can wait. On 3.8 Flash that's $0.375 in and $1.875 out until December 31.
  • Priority. 1.8 times Standard on the Flash models ($1.35 / $6.75 on 3.8 Flash) for requests that can't be slowed down.
  • Grounding with Google Search. On Gemini 3 models, 5,000 search requests a month are free, shared across all of them, then $14 per 1,000. One user request can trigger several searches, and Google bills each one.
  • Long prompts on Pro. Above 200,000 tokens, 3.1 Pro Preview goes from $2 / $12 to $4 / $18.
  • Image and video. Priced separately. One deadline already passed: Google's page says Gemini 2.5 Flash Image was shut down on October 2, 2026, so anything still calling it needs moving to 3.1 Flash Image.

What I'd do this week

If you already call Gemini, look up which model each feature uses. Anything on 3.5 Flash moves to 3.8 Flash today. Then put January's price, not today's, in next year's budget. If you're choosing a vendor, run your real prompts through two models from the cost table above for a week and compare quality per dollar, not price per token.

Sources

Free Resource Justin McKelvey

Hit the cap? Every plan's limit, one page.

ChatGPT, Codex, Claude, Gemini, and Copilot: what's metered, what resets when, and the cheapest way around each. Re-sent when a limit moves.

Frequently Asked Questions

How much does the Gemini API cost?
As of October 4, 2026, Google's rate card lists Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, rising to $1.50 and $7.50 on January 1, 2027. Gemini 3.1 Pro Preview is $2 in and $12 out for prompts up to 200,000 tokens. The cheapest text model is Gemini 2.5 Flash-Lite at $0.10 in and $0.40 out. Batch requests cost half.
Is the Gemini API free?
There is a free tier with limited access to certain models, and Google AI Studio itself is free to use. The catch is in Google's own table: on the free tier, your content is used to improve Google's products. On the paid tier it isn't. Gemini 3.1 Pro Preview has no free tier at all.
Is Gemini API pricing going up?
Yes, for the newest Flash models. Google's pricing page lists Gemini 3.8, 3.7 and 3.6 Flash at $0.75 input and $3.75 output through December 31, 2026, and $1.50 and $7.50 starting January 1, 2027. Batch, Flex, Priority and caching rates double on the same date.
Is Gemini cheaper than Claude or OpenAI's API?
At the Flash tier, yes. On one workload of 150 million input and 30 million output tokens a month, Gemini 3.8 Flash costs $225 now and $450 from January, against $600 for Claude Sonnet 5.5 or GPT-6.1 Sol and $300 for Claude Haiku 4.5. At the very bottom, GPT-6 Luna ($30) and Gemini 2.5 Flash-Lite ($27) are close to free.
How much does Google Search grounding cost on the Gemini API?
On Gemini 3 models, the first 5,000 grounded search requests a month are free, shared across all Gemini 3 models, then $14 per 1,000. One user request can trigger several searches, and each one is billed.

More on AI for Business

AI for Churches (2026): What It's Good For, Where to Keep It Out, and the Discount Paperwork Nobody Mentions

AI can take the bulletin, the volunteer emails and the sermon clips off a church office's plate for about $64 a month for eight staff. The catch is the nonprofit discount: the IRS never required most churches to apply for exemption, and the AI vendors' verifier wants proof. Here's what to use it for, what to keep out, and how to get verified.

7 min

ChatGPT Business vs Plus (2026): Which One a Small Business Should Actually Pay For

ChatGPT Plus vs Business, as of September 2026: Plus is $20 a month for one person and OpenAI may train on your chats unless you opt out; Business is $20 a seat billed annually or $25 monthly, two seats minimum, with no training by default, admin controls, shared projects for groups, and the Company Knowledge plugin. The side-by-side from OpenAI's pages, and three owner scenarios with a verdict each.

6 min

ChatGPT for Customer Service (2026): What Works, What Breaks, and What It Costs a Small Business

ChatGPT for customer service, as of September 2026: it drafts replies, answers policy questions from your own documents, and triages an inbox for $20 to $25 a seat a month on ChatGPT Business. It can't look up an order, issue a refund, or remember a customer unless you connect the system that holds them. The honest capability line, the plan you need, a setup in six steps, the failure modes, and the point where a real agent takes over.

7 min

ChatGPT Pulse (2026): What It Was, Why OpenAI Retired It, and How to Get the Morning Brief for Your Business

ChatGPT Pulse, the daily research cards OpenAI previewed for Pro users in September 2025, was sunset on June 17, 2026 with a 14-day wind-down. What it did, which plans ever had it, why scheduled tasks replaced it, and how an owner rebuilds the morning brief on a Plus or Business seat without handing ChatGPT the wrong things.

7 min
Justin McKelvey, Fractional CTO and AI consultant in Austin, TX

Written by

Justin McKelvey

Fractional CTO & AI consultant in Austin, TX. 15 years building software, 50+ products shipped, $53M+ in client revenue generated. I help $1M–$50M founders ship production software and automate operations with AI, without hiring a full-time executive team.

Work with me

If this was useful, here are two ways I can help: