ATAILA Newsroom · Budapest · 2026-09-04
2026-W36
Public AI is not getting cheaper. Three announcements in one quarter say so.
GitHub has moved Copilot from flat premium-request units to usage-based billing on token consumption, and said why in a sentence worth keeping: “GitHub has absorbed much of the escalating inference cost behind that usage, but the current premium request model is no longer sustainable.” That is a vendor stating on the record that it was subsidising the product and has stopped. It is not the only one this quarter.
The article we are responding to
„GitHub Copilot is moving to usage-based billing”
The GitHub Blog · 2026-04-27
What is actually happening
Three separate announcements, one quarter, one direction:
1 · GitHub
Flat rate becomes metered
Copilot moved from premium request units to billing on token consumption — input, output and cached — at each model's API rates. Base subscription prices did not change; what those prices include did. GitHub's stated reason is that agentic use has far higher compute demand, and that a quick question and a multi-hour autonomous session used to cost the customer the same.
2 · Anthropic
A temporary allowance expires
From 13 May a promotion gave Claude Code subscribers 50% higher weekly limits. It ends on 13 September, when limits return to standard levels. To be precise: this is a bonus expiring, not a cut below the baseline, and no price changes. The practical effect is still that the same subscription buys less work next week than it bought last week.
3 · Google
Same unit price, bigger bill
Gemini Flash 3.8 is priced identically to its predecessor per token — and early Artificial Analysis measurements put the cost per task up to 40% higher, because it reasons in more steps. Nobody raised a price; the invoice moves anyway. We wrote about this separately.
Why it is structural
It would be comfortable to read these as three unrelated commercial decisions. They are not. They are what the end of a subsidised phase looks like from the customer side, and the arithmetic underneath is public.
PwC's datacentre outlook, reported by economx on 2 September, puts roughly $800 billion of datacentre investment in 2026 alone, rising toward $1.8 trillion a year by 2050, with the cumulative figure approaching $31.6 trillion. It also notes the detail that decides the economics: the useful life of the GPUs at the centre of it is four to six years.
That is the whole story. Capital on that scale, depreciating that fast, has to be recovered from someone — and the someone is whoever is renting the inference. A price that was set to win the market is not the price that services the debt behind it.
None of that makes the vendors villains. GitHub said the quiet part out loud and did the honest thing: it stopped pretending an agentic session costs what a chat message costs. But an EU company budgeting for next year should stop assuming the trend line points down. Every signal this quarter says it points up.
What we think — including where our own pitch is too easy
The tempting version of this argument is: buy your own hardware and you only pay for electricity. We are not going to write that, because our own site argues against it. Buying a box gets you the box. The patching, the backups, the failing GPU at three in the morning, the model lifecycle and the depreciation PwC just put a number on are all still yours, and they are the expensive part.
The honest claim is narrower:
Owned capacity changes what your cost depends on
Metered public AI makes your run-rate a function of someone else's pricing decisions, release schedule and funding pressure. Dedicated capacity — ours or yours — makes it a function of your own workload. It is not automatically cheaper on day one. It is forecastable, which is what a finance function actually needs, and it stops moving when a vendor's board decides the subsidy is over.
The stack is the work, not the hardware
If you do own hardware, the electricity really is close to the marginal cost of the next inference. But getting from "we own GPUs" to "a confidential workflow runs in production with backups, monitoring, a release path and an SLA" is a platform — and building that yourself is the project that eats the savings. That is the part we operate.
The cheapest AI in 2027 will be the one you did not have to re-budget
We price one operated workflow at a flat monthly fee, with no per-token line. When inference costs move, that is our problem to absorb, not a surprise on your invoice.
We are not claiming private AI is always cheaper than a public API — for spiky, low-stakes work it usually is not, and we will say so. We are claiming that if a workflow matters enough to run every day, you should know what it will cost next year. On metered public AI, nobody can honestly tell you that right now.
Where we stand
We use public models ourselves where the data allows it, and we will keep doing so. This is not a call to move everything in-house.
It is a call to notice the direction. Three vendors in one quarter have either started charging for what used to be included, or stopped including as much. The capital behind the buildout says that continues. If your plan for next year assumes AI gets cheaper per unit of work, that plan now needs a second column.
The original article
„GitHub Copilot is moving to usage-based billing”
The GitHub Blog · 2026-04-27
← Back to the Newsroom Press inquiries: contact us