Saturday, 12 September 2026
ZAR/USDR16.160.06%. Rand weaker against the US dollar
ZAR/EURR18.730.13%. Rand stronger against the euro
ZAR/GBPR21.830.00%. Rand flat against the pound
Tech & Telco

Google’s cheap AI tier doubles in price on 1 January. If your business runs on it, budget now.

Google’s cheap AI tier doubles in price on 1 January. If your business runs on it, budget now.

A price that is scheduled to double is a great deal easier to plan around than a price that merely might. Google has published both halves of that arithmetic for Gemini 3.8 Flash, and the second half takes effect on 1 January 2027.

Gemini 3.8 Flash shipped on 2 September 2026. Its published rate is $0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026. On 1 January 2027 both figures double, to $1.50 and $7.50. Those numbers and that date are carried consistently across independent reviews of the model, including eesel AI, Layer3Labs and CellCog, which agree on the rates and on the switchover.

At the rand rate shown on this site’s own market data at the time of writing, R16.15 to the dollar, the input rate moves from roughly R12.11 to about R24.23 per million tokens, and output from around R60.56 to R121.13. Currency moves, so treat the rand figures as a conversion at today’s rate rather than a forecast. The doubling itself does not move.

Why this particular tier is the one smaller businesses should care about

A token is the unit these models are billed in, roughly a fragment of a word, so a page of text runs to several hundred tokens and a long document to tens of thousands. Input tokens are what you send the model. Output tokens are what it sends back, and they cost about five times as much, which is why a system that writes long answers costs far more to run than one that classifies or extracts.

The Flash line is the cheap, fast tier, and it is where most real business workloads actually sit. The frontier models get the headlines, but the things companies quietly run all day, sorting support email, pulling fields out of invoices, drafting first-pass replies, summarising documents, checking submissions against a rule, do not need a frontier model and are not usually built on one. They are built on whatever is cheapest per unit of adequate output. That makes a change in the cheap tier’s list price a change in the unit economics of a lot of small South African software.

The second increase that will not show up as a price change

There is a compounding effect underneath the headline doubling, and it is easy to miss because it never appears as a rate.

Gemini 3.8 Flash is built on its predecessor rather than on a new base model, and it improves partly by working harder: spending more reasoning effort, and therefore more tokens, on the same question. Those reasoning tokens are billed at the output rate and the customer cannot see them. CellCog makes the point directly, noting that identical per-token rates can still produce a higher bill for equivalent work.

Put the two together and a business that budgets its 2027 AI costs by doubling a September invoice will be wrong in the same direction twice. The rate doubles, and the number of billable tokens consumed per task may already be higher than it was on the previous model.

What to do with the window that is left

Nobody is negotiating a published list price, so the only lever available is consumption, and the useful work is measurement rather than panic.

The first thing worth knowing is what you actually spend per unit of business value, not per month. Cost per processed invoice, per resolved support ticket, per generated summary. A monthly total tells you nothing about whether the January change breaks the model, because it does not separate volume growth from unit cost. A per-task figure does, and doubling it is a five-second forecast.

The second is where the output tokens are going. Output is the expensive half, and a surprising share of it is often waste: a model asked an open question when a constrained one would do, or returning prose when the system only needs a field. Tightening what you ask for reduces the bill under either price.

The third is portability. A workload that can only run on one provider’s model has no response available when that provider changes its pricing. One that has been tested against an alternative has a negotiating position, or at least an exit. The four months to 31 December is a reasonable amount of time to find out which kind you have.

The wider picture on AI pricing

The comfortable assumption in the market has been that the cost of AI only falls. It has mostly fallen, and this site reported on 10 September that Anthropic cut its cache-token price sharply with Claude Fable 5.1, which is real and points the other way.

What the Gemini schedule shows is that the direction is not uniform and is not guaranteed. Introductory pricing is a commercial decision, and the point of an introductory rate is to end. A business that treated the low rate as the permanent cost of a capability, and priced its own product accordingly, has made a planning assumption rather than an observation.

Google also shipped a separate security-focused version of the same model, Gemini 3.8 Flash Cyber, which is not on the price list at all because it is not generally for sale. That is a different story, and a more uncomfortable one for South African buyers.

For now the arithmetic is unusually kind: the change is published, dated and exactly a factor of two. There are not many cost increases a business gets to see coming this clearly, and the ones it does see coming are the ones there is no excuse for being surprised by.

Gemini 3.8 Flash is the second frontier AI story this site has covered in a week from the cost side, alongside Claude Fable 5.1’s own pricing changes. Read together with GPT-6 Astra’s cybersecurity threshold and South Africa’s own withdrawn AI policy, the pattern is the same one running through all three: adoption is outrunning both regulation and budget planning. For a starting point on managing the cost side specifically, see our practical guide to AI tools for small business.

This report is based on a statement available at www.eesel.ai.