(434) 236-9027

Two Frontier Models Launched Three Days Apart at the Same Price. Memory Is Now a Third of a Phone.

A single green memory module standing upright on a dark wood desk in low morning light, with an out of focus stack of paper behind it
Bottom line

Anthropic and OpenAI shipped new flagship models three days apart in early September 2026, and both landed on exactly the same headline price: $10 per million input tokens, $50 per million output tokens, $12.50 per million for cache writes. The number that will actually decide your bill is the one they did not match. Fable 5.1 reads from cache at $0.25 per million and charges nothing extra for a long prompt. GPT-6 Astra reads at $1.00 and adds a surcharge above 272,000 tokens. If you ask one off questions, none of this matters. If you run agents in a loop, it is the line that separates the two bills. Meanwhile memory has grown to about a third of what an iPhone costs to build, which matters before you approve a hardware refresh.

Two labs, one price

In early September 2026 two frontier models arrived three days apart carrying identical headline pricing. Anthropic released Claude Fable 5.1 at $10 per million input tokens and $50 per million output tokens. OpenAI released GPT-6 Astra at the same $10 and $50. Both charge $12.50 per million for cache writes.

They got there by opposite routes. OpenAI's previous model, GPT-5.6 Sol, charged $4 input and $20 output, so Astra is exactly 2.5 times the price of the thing it replaces. Anthropic did not raise anything. Claude Fable 5 was already $10 and $50, and 5.1 held that completely flat across the version bump, putting the entire change into one line item by dropping cache reads from $1.00 per million to $0.25.

Anthropic's own published benchmarks for 5.1 report a Terminal-Bench-Science 0.1 score of 52.6 percent against Fable 5's 24.7 percent, Humanity's Last Exam at 65.0 percent with tools and 60.9 percent without, and Terminal-Bench 4.0 at 55.8 percent. Those are the vendor's numbers on the vendor's page, which is worth saying plainly rather than repeating them as though someone independent produced them.

Published list prices, per million tokens, checked September 3, 2026
ModelInputOutputCache writeCache read
Claude Fable 5.1$10.00$50.00$12.50$0.25
GPT-6 Astra$10.00$50.00$12.50$1.00
Claude Opus 5$5.00$25.00$6.25$0.50
GPT-5.6 Sol$4.00$20.00not listed$0.40
Claude Sonnet 5$2.00$10.00$2.50$0.20
Gemini 3.8 Flash$0.75$3.75not listednot listed

Gemini 3.8 Flash pricing is promotional through December 31, 2026. GPT-6 Astra prices shown are the short context tier, below 272,000 tokens.

The number that decides your bill

Cache reads are where the two models genuinely part company, and it is the line most people skim past because it has the smallest dollar figure attached to it.

Ask a model one question and its context gets read once. Point an agent at a codebase or a document set and tell it to work through a task, and it reads much of that same context again on every turn, sometimes dozens or hundreds of times before it is done. For the part of the context that does not change, you write once and read many times, which is why in a long running session the cache read price stops being a rounding error and becomes the recurring cost.

That is a floor rather than the whole bill. A context that grows every turn writes its new tail each time, and a cache that expires has to be written again, so real sessions pay more writes than the clean version of this story suggests.

Fable 5.1 reads at $0.25 per million tokens, which is 0.025 times its base input price. Every other Anthropic model uses a 0.1 multiplier, and Fable 5.1 and Mythos 5.1 are the only two off it. GPT-6 Astra reads cached input at $1.00 per million, a standard 0.1 of its base, so the operation an agent performs most often costs four times as much there.

Long context is priced differently by each lab

The second divergence compounds the first. Anthropic's documentation states that Claude 4.6 and later models include the full one million token context window at standard pricing, so a 900,000 token request bills at the same per token rate as a 9,000 token one. GPT-6 Astra has a context window of 1,050,000 tokens and a maximum output of 128,000, but above 272,000 tokens it charges 2x input, 2x cache reads and 1.5x output. That works out to $20 input and $75 output on the long tier.

Anthropic is pricing long context as the normal operating condition. OpenAI is pricing it as a premium capability. They produce very different invoices for the same job, and the gap falls hardest on an agent that re-reads a large body of material on every turn.

The two differences stack. Fable 5.1 holds its cache read rate steady no matter how long the prompt gets, while Astra's doubles to $2.00 per million once the threshold is crossed, which opens an eightfold gap on cache reads alone before the doubled input rate is counted at all. Whether that matters to you depends entirely on whether your prompts are long and whether you re-read them.

The question to ask your own stack

Pull one representative agent session and check two numbers: how many tokens sit in the context, and how many times that context is read before the session ends. Those two figures decide which of these vendors is cheaper for you. Nothing about the headline price tells you.

The cheap tier has a date on it

Google's Gemini 3.8 Flash undercuts both frontier models by roughly thirteen times on input, at $0.75 per million input and $3.75 per million output. Read the footnote though. Those prices run through December 31, 2026. On January 1, 2027 they become $1.50 and $7.50, which is a doubling on a date that is already published.

Anyone sizing a bulk workload on $0.75 input should size it on $1.50 as well, since both numbers are printed on the same page and only one of them is temporary.

Dated prices do move in both directions, which is the argument for putting a calendar reminder on the pricing page rather than trusting your notes. Claude Sonnet 5 launched at $2 input and $10 output as introductory pricing, scheduled to rise to $3 and $15 on September 1, 2026. That increase was cancelled and the lower price became the standard one. Anyone who re-read the page found a better number than the one in their forecast. Anyone who assumed the rise happened is now budgeting against a price that does not exist.

Memory became the expensive part

The hardware side of the same month is less discussed and lands closer to home.

TrendForce reported on August 10, 2026 that the bill of materials for a 256GB iPhone 18 Pro is running about 38 percent higher than its 2025 predecessor, and that memory has gone from roughly 10 percent of that bill a year ago to about 34 percent in the third quarter of 2026, with a forecast above 40 percent in the first half of 2027. Its September 3 follow up put memory cost for that model nearly 400 percent higher than a year earlier and expects a 10 to 20 percent retail price increase, describing memory prices as having entered a major upcycle in the second half of 2025.

Worth being careful about the explanation. It is tempting to draw a straight line from AI datacenter buildout to the price of the phone in your pocket, and that story may well be right, but TrendForce does not make that claim in either release. What it documents is the upcycle and the cost, not the cause. TrendForce also expects Apple to absorb some of it by accepting lower gross margin, following what it did on recent MacBook launches, so the retail increase should land softer than the component increase.

What this changes for a refresh

On the one device TrendForce broke down, memory is about a third of what it costs to build, and the forecast has that share rising into 2027. If you have a hardware refresh already budgeted, the usual instinct to wait for a better price does not have data behind it right now.

Apple is selling memory as the AI feature

Apple announced the M6 and M5 Ultra on August 25, 2026. The M6 is Apple's first 2 nanometer chip, and it ships in the Mac mini with a Dual 16 core Neural Engine, up to 170GB/s of unified memory bandwidth and up to 32GB of RAM. Apple claims the Neural Engine change delivers up to twice the peak compute of previous generations.

The M5 Ultra is the interesting one for anyone running models locally. It ships in the Mac Studio, joining two dual die M5 Max chips into a quad die design, and reaches 1.2TB/s of unified memory bandwidth with up to 512GB of unified memory. Apple's claim is up to 4.5 times the peak GPU compute for AI compared to M3 Ultra and over six times the M1 Ultra.

Those performance multiples are Apple's, measured by Apple. The product decision underneath them is the tell: the specification Apple leads with for AI work is memory capacity and bandwidth, while memory climbed to about a third of what an iPhone costs to build.

How I actually route this

I run a mixed stack. There is a flat fee subscription covering about a dozen models, and separately there is metered API access to the frontier models. Those two things want opposite treatment. The subscription is a sunk cost, so the only mistake is underusing it. The metered access is the opposite, and every token there should be buying something the cheap tier could not produce.

That produces one rule I apply without much thought now. Flat fee and free models do drafting and bulk work: first passes, first drafts, summarising, the mechanical parts of a review. Metered frontier models do verification, judgment, and anything a client will read. The metered bill stays small because by the time it is spending anything, most of the work is already done.

The second rule took longer to learn. Do not verify a model's output using another model from the same family, because same family models tend to agree with each other, and an agreement you were always going to get is not a check. Something drafted by one architecture and reviewed by a genuinely different one catches errors that no amount of self review surfaces. Most of the real mistakes caught in my own work came from that pairing rather than from any single model being clever.

Here is the cache read arithmetic on a single agent session holding a 100,000 token context and reading it fifty times.

Cache read cost, one session, 100,000 token context read 50 times
Model and tierRate per millionCost for 5M token reads
Claude Fable 5.1, any prompt length$0.25$1.25
GPT-6 Astra, under 272,000 tokens$1.00$5.00
GPT-6 Astra, same reads if the request sits above 272,000 tokens$2.00$10.00

Cache reads only. Input, output and cache writes are billed on top, and are identical between the two models below Astra's threshold.

At a handful of sessions a week nobody notices any of this. At a hundred a week it quietly decides whether a workflow gets to run continuously or has to be rationed, which is usually the moment people go looking for the pricing page in the first place.

On hardware, my honest position is softer. Memory prices are rising rather than falling, the forecast has that continuing through the first half of 2027, and TrendForce expects vendors to eat part of it. If a refresh is already in the budget, I would not hold it back waiting for relief that nothing in the current data predicts. That is a read on a forecast, not a fact, and I would rather say so than dress it up.

The expensive part moved, on the API bill and inside the phone, and neither vendor put it in the headline.

Mr. Botsworth

Mr. Botsworth

Hey! I'm Mr. Botsworth, Greg's search bot. Ask me about his projects, skills, or services.