Claude Sonnet 5 pricing: the September 1 rise is bigger than it looks
There are twenty five days left on Claude Sonnet 5’s introductory pricing, and most of what I have read frames what happens on 1 September as a fifty per cent increase. That number is accurate, and it is also the least useful thing you can know about the change, because it describes the price list rather than the bill. Sonnet 5 counts tokens differently from the model most teams are migrating away from, and once you hold those two facts together the picture moves quite a long way.
Here is the short version. If you are running Sonnet 5 today, the same workload costs fifty per cent more on 1 September. If you are still on Sonnet 4.6 and you migrate, the same work costs about thirty per cent more than it does now, even though the published price you move onto is identical to the one you are already paying. And from 1 September, Sonnet 4.6 and Sonnet 5 sit on the same page at exactly the same price per million tokens while one of them is roughly twenty three per cent cheaper to actually run.
None of that requires inside information. It follows from two numbers Anthropic publishes itself, and the only reason it is not the headline is that the two numbers live on different pages.
What actually changes on 1 September
The pricing page is unambiguous about the first number:
“Introductory pricing of $2/$10 per million input/output tokens is in effect through August 31, 2026, after which the standard pricing of $3/$15 per million input/output tokens will take effect.”
So Sonnet 5 moves from $2 and $10 to $3 and $15 per million input and output tokens. That is fifty per cent on both dimensions, and it is the figure being quoted everywhere. Worth noting is where it lands: $3 and $15 is exactly what Sonnet 4.6, Sonnet 4.5 and Sonnet 4 have always cost. Sonnet 5 is not becoming expensive by Sonnet standards, it is simply arriving at the standing Sonnet price after a discount window.
If the token were a fixed unit, that would be the whole story, and the correct advice would be to budget fifty per cent more and move on.
The part that is not on the price list
Further down the same pricing page, in a note under the model table:
“Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Claude Sonnet 4.6 and earlier models use the previous tokenizer.”
Read that against the model list and the consequence is specific. Sonnet 5 is a 4.7-or-later model, so it uses the new tokenizer. Sonnet 4.6 is explicitly named as using the old one. The same prompt, the same document, the same repository sent to both models produces roughly thirty per cent more billable tokens on Sonnet 5 than on Sonnet 4.6, before anyone changes a line of code.
Anthropic’s migration guide says the same thing in the other direction, which is a useful corroboration rather than a second source: “The same input text produces approximately 30% more tokens than on Sonnet 4.6.” Per-token pricing is unchanged by the tokenizer. The cost of a request is not.
To be fair to the coverage, a handful of write ups have flagged the tokenizer. What I have not seen anyone do is put it through the arithmetic alongside the date, which is where it stops being trivia and starts being a budget line.
One caveat before any of the numbers, because it is the number everything else rests on. Thirty per cent is a central estimate rather than a constant, and Anthropic says so in the same breath: “the exact increase depends on the content and workload shape.” Across content types the observed inflation runs from roughly no change at all up to about thirty five per cent, and dense material such as code and structured data tends to sit at the top of that band while ordinary English sits lower. So treat every figure below as a reading at 1.3x, and measure your own before you budget against it. I have made the tokenizer figure the first argument to the script at the end of this post for exactly that reason.
One part of this is not sensitive to that at all, and it is worth separating out. If you are already on Sonnet 5, your bill rises by fifty per cent on 1 September whatever your workload does with tokens, because that leg is a pure price change and the tokenizer applies equally on both sides of the date. It is only the comparison against Sonnet 4.6 that moves with the inflation figure.
Doing the arithmetic
I priced one unit of identical work, meaning the same prompts and the same outputs, across the three states a team can be in. I used a 75/25 split between input and output tokens, which is typical for agentic and coding workloads where a large context produces a comparatively small edit. The split changes the exact figures a little and changes none of the conclusions, because both dimensions rise by the same fifty per cent and the tokenizer applies to both.

Three things fall out of that, and the middle one is the one I would put in front of a finance team.
Today, Sonnet 5 is about thirteen per cent cheaper than Sonnet 4.6 for the same work. The introductory discount is large enough to swallow the tokenizer and leave change. Teams that migrated during the window and measured their spend will have seen a real saving, which is exactly why the next point catches people.
On 1 September, that same workload costs fifty per cent more than it does today and about thirty per cent more than it did on Sonnet 4.6. The fifty per cent is the price change. The thirty per cent is what you are actually comparing against if you migrated from 4.6, and it is the number that belongs in a forecast.
The rise arrives without a deploy. No model change, no config change, nothing in your repository moves. Most cost dashboards report dollars rather than tokens, so a team that has not been told about the date will see a step change in spend and start looking for a traffic explanation that does not exist.
Two models, one price, two costs
This is the part I find genuinely interesting, and the reason I think the story is worth more than a budget warning.
From 1 September, Sonnet 4.6 and Sonnet 5 are both listed at $3 per million input tokens and $15 per million output tokens. Identical rows on an identical page. But because Sonnet 5 emits roughly thirty per cent more tokens for the same text, running your workload on Sonnet 4.6 costs about twenty three per cent less than running it on Sonnet 5, at a price the page says is the same.

I want to be careful about what that does and does not mean, because it would be very easy to read it as advice to downgrade, and it is not. Twenty three per cent cheaper to run is a statement about cost per token, not cost per result. Sonnet 5 is a materially more capable model than Sonnet 4.6, particularly on coding and agentic work, and on the tasks where that capability shows up it will often reach the answer in fewer attempts, with fewer retries and less human correction. A model that costs twenty three per cent more per token and gets there first try is cheaper per finished task, and the finished task is what you are actually buying.
So the honest framing is narrower than the headline. If Sonnet 4.6 already does your job to your standard, it is now the cheaper way to keep doing it, and that is worth knowing. If your workload is one where Sonnet 5’s gains are real, pay for them, and use the effort lever below to take the edge off the difference. What you should not do is pick between them on a per-token price, in either direction.
The wider consequence is that a per-token price is no longer a figure you can compare across Claude generations, and it never was one you could compare across vendors. Any spreadsheet that ranks models by dollars per million tokens is now comparing quantities measured in different units, which is a category error rather than a rounding error. The honest unit is cost per completed task on your own workload, and the only way to get it is to measure.
The mitigation that is sitting in Anthropic’s own docs
It would be easy to stop there and file this as a price rise, but that would leave out the half of the picture that makes Sonnet 5 worth migrating to in the first place, and it would make this post less useful than it should be.
Anthropic’s migration guide offers a mapping between the two models at different effort levels:
“Claude Sonnet 5 at medium is comparable in intelligence to Sonnet 4.6 at high, and Claude Sonnet 5 at high is comparable to Sonnet 4.6 at max.”
That changes the calculation for a lot of workloads. Effort governs how much the model thinks and how many tokens it spends getting to an answer, so dropping one level is a direct reduction in output tokens. If your workload currently runs Sonnet 4.6 at high and you can move to Sonnet 5 at medium for equivalent quality, you are recovering a meaningful share of the thirty per cent, and you are getting a stronger model on coding and agentic work while you do it. Whether it is all of the thirty per cent, or half, depends entirely on your own traces, which is why I am not going to publish a number I have not measured on a real workload.
Two other levers are unchanged and are proportionally more valuable now than they were a month ago. The Batch API still takes fifty per cent off both input and output for work that does not need to be synchronous, and prompt caching still serves cached reads at a tenth of the input price. Neither is new, and both are worth more when the base rate goes up.
What I would do in the next twenty five days
This is the part that survives the news cycle, so it is the part I would act on.
Measure your own tokenizer delta rather than applying thirty per cent. Run
count_tokensagainst Sonnet 5 on a representative sample of your real prompts and compare it to the same call against Sonnet 4.6. Anthropic is explicit that the figure varies by content and workload shape, and code and structured data are exactly the kind of dense content where a tokenizer change bites hardest. A blanket multiplier will be wrong in one direction or the other.Check your output ceilings before the bill. A
max_tokensvalue tuned against Sonnet 4.6 can truncate equivalent output on Sonnet 5, because the same answer is now more tokens. That shows up as a quality regression rather than a cost one, which makes it harder to attribute. Compaction thresholds have the same problem.Sweep effort down one level and judge it on your own evals. This is the single biggest lever available and it costs nothing but an afternoon. Anthropic’s own mapping suggests the headroom exists; only your evals can tell you whether it exists for your workload.
If you are happy on Sonnet 4.6, know that it is not going anywhere. It is not deprecated and it is not retired, and from 1 September it is the cheaper of the two per unit of work. That is a legitimate position to hold if your workload is not one where Sonnet 5’s gains show up.
Tell whoever owns the budget the date. A fifty per cent step change on a fixed calendar day, with no deploy attached to it, is the kind of thing that costs a team an afternoon of incident review if nobody saw it coming.
The thing worth remembering after the date passes
The specific number stops mattering on 2 September. What does not stop mattering is that a token quietly became a per-generation unit rather than a stable one, and the industry’s default way of comparing model costs did not notice.
We have spent two years treating dollars per million tokens as though it were a price per litre, comparable across vendors and across versions. It is closer to a price per box, where the boxes are different sizes and the size is documented in a footnote. Anthropic has been upfront about this, and the note is right there on the pricing page, but it sits below the table that everybody screenshots. The benchmark posts comparing cost per task across models are affected by the same thing, and most of them were computed with token counts that are no longer directly comparable to the ones they are being compared against.
The correction is not complicated. Price your own workload, on your own traces, in dollars per completed task, and treat any per-token figure as a component of that rather than a substitute for it. That was always the right method. The tokenizer change just removed the option of pretending otherwise.
Sources
Anthropic: Pricing. The model table, the introductory pricing note with the 31 August date, and the tokenizer note quoted above.
Anthropic: Model migration guide. The Sonnet 4.6 to Sonnet 5 section, covering the tokenizer delta and the effort-level mapping.
Anthropic: Introducing Claude Sonnet 5. The launch announcement.
The Register: Anthropic’s extravagant tokenizer complicates AI pricing. Earlier coverage of the same tokenizer change as it applied to the Opus line.
The arithmetic in this post is reproducible from the published rates and Anthropic’s own figure. I have put it in a short script in our repository that takes the tokenizer inflation as its first argument and prints a sensitivity table, so you can see what the conclusions do at 1.1x or 1.35x rather than taking 1.3x on trust. Figures are correct as of 7 August 2026 and describe first-party Claude API pricing; Amazon Bedrock and Google Cloud are priced separately by those providers.
See it instead of reading about it.
Codus is out now on macOS.