ATAILA Newsroom · Budapest · 2026-09-04
2026-W36
The price per token did not move. The bill can rise up to 40%.
On 3 September HWSW reported something quietly important about Google's Gemini Flash 3.8: the published price did not change at all — the same $0.75 per million input tokens and $3.75 per million output tokens as the previous version. But early Artificial Analysis measurements suggest the cost per task can be up to 40% higher, because the new model takes more reasoning steps on complex work. Nobody raised a price. Everybody's bill can still go up.
The article we are responding to
„Keményebben dolgozik a Google új modellje, de ennek ára van”
HWSW · 2026-09-03
What the article reports
The model is, by the article's account, genuinely better at software development and agent work — better than comparable options at the same headline price. This is not a story about a vendor doing something underhanded. It is a story about what a per-token price list can and cannot tell you.
What we think
A unit price is only a budget if you know how many units you will consume. With reasoning models you do not, and increasingly you cannot: consumption is a property of the model's behaviour, and that behaviour changes when the vendor ships an upgrade you did not choose and cannot decline.
The upgrade you did not ask for changes your costs
Nobody at Google sent an invoice increase. A better model arrived, it thinks in more steps, and the same workflow now costs more to run. If your finance team approved a budget on the old numbers, that budget is wrong and nobody told them.
"Cheaper per token" is not "cheaper"
Per-token comparisons between models are close to meaningless once step counts differ. The only number that matters is what one completed unit of your work costs — a processed invoice, a reviewed contract, an answered query — and that is not on anyone's pricing page.
This is a forecasting problem, not a pricing problem
We have written before that most enterprise AI dies between the pilot and production. Metered inference is one of the reasons. A pilot that costs a rounding error becomes a line item nobody can defend at 100× the volume, and the CFO's question — what will this cost next year? — has no honest answer.
We price one operated confidential workflow at a flat monthly fee. Never per user, never per token — your invoice does not carry a usage number. When we ship a better model, your bill does not move. That is the point of the pricing, not a promotion.
Where we stand
We are not against metered pricing everywhere. For spiky, experimental, low-stakes work it is exactly right, and we use public models ourselves where the data allows it. The problem is what happens when a metered dependency ends up underneath something a business runs on every day. Then the bill becomes a function of a vendor's release schedule, and you are the one absorbing the variance.
If you are budgeting AI for next year, run the test this article implies: take your busiest workflow, and ask what it costs to run today, and what it will cost if your provider ships a model that reasons twice as hard. If you cannot answer the second question, you do not have a forecast — you have last month's invoice and some hope.
The original article
„Keményebben dolgozik a Google új modellje, de ennek ára van”
HWSW · 2026-09-03
← Back to the Newsroom Press inquiries: contact us