What actually changed
Two things moved. One is a price cut hidden below the headline, and the other is a setting that finally does what it claims.
Anthropic released Claude Fable 5.1 on 1 September 2026. The coverage led with benchmark scores, and the scores are real. On one test of long research work it scored 52.6% where the previous model scored 24.7%. That is the sort of number that gets reposted and then changes nothing about your week.
The advertised prices did not move. Sending text in still costs 10 dollars per million words-worth of text, and getting text back still costs 50. What moved is a line called the cache read, which fell from a dollar to 25 cents. That sounds like a rounding error and it is the whole story, for reasons the next section explains.
The second change is a setting called effort. It has been there for a while and, on the old model, turning it up mostly wasted money. On this one it works. That means the cheapest setting is now genuinely usable for a lot of real work, which is where most of the saving actually comes from.
One honest limit before you read on. Everything here is measured on someone else's tasks, mostly Anthropic's own and one independent set from Cursor. None of it tells you how it behaves on your documents, your writing or your customers. The twenty-minute test further down is the only part that does.
The four words you need
- Token
- A chunk of text, roughly three quarters of a word. AI is billed per million of them. When you see a price like 10 dollars per million tokens, read it as about 10 dollars per 750,000 words.
- Cache read
- Re-reading something it has already been shown. Ask it to work through a 200-page contract and it does not read the contract once. It re-reads the relevant parts over and over as it works, and you were being charged full price every time.
- Effort
- How long it thinks before answering. Five levels: low, medium, high, extra high, maximum. More thinking costs more and takes longer. It does not always produce a better answer, which is the point of this guide.
- Agent
- A job you set going and walk away from, rather than a conversation you sit through. It works through steps on its own and tells you when it is done. This is the kind of work that got dramatically cheaper.
Why 25 cents matters more than 10 dollars
On a short question the cache barely registers. On a long job it is most of what you pay.
Ask a quick question and you send a bit of text and get a bit back. You are paying the headline prices and the cache hardly comes into it.
Now give it a full contract and ask for every clause that creates a liability. It does not read the document once and answer. It works through it, holds it in mind, checks back, cross-references, and re-reads the same pages repeatedly. Every one of those re-reads was charged at the old rate.
That is why a cut to one line produces a much larger cut to the bill than the line's size suggests. Anthropic's own estimate is around 25% off ordinary work and up to about 45% off the long unattended jobs, and the longer the job runs the closer you get to the larger number.
The same job, two settings
Cursor, a company that sells AI coding tools, published its own measurements rather than relying on Anthropic's. Two rows from that table, running the same set of tasks.
Previous model, maximum effort
Scored 70.5%. Used 103,525 tokens per task and cost 17 dollars 32.
Fable 5.1, medium effort
Scored 68.0%. Used 23,801 tokens per task and cost 3 dollars 53.
Why it changed. Two and a half percentage points separate the answers. Roughly four times the tokens and nearly five times the money separate the bills. On most business work, two and a half points is not a difference you would notice in the output, and the cost gap is one you would notice on an invoice.
Why turning it up used to be a waste
On the previous model, paying for more thinking bought you almost nothing above the middle setting.
Anthropic published a chart plotting score against cost for both models across all five effort levels. On the previous model the line climbs to the middle and then stops. High effort scored 25.2%. Extra high, which costs more, scored 23.6%. Maximum recovered slightly to 24.8%. You could roughly double the bill and end up level, or slightly behind.
On Fable 5.1 the line keeps climbing the whole way: 26.4, then 35.8, then 40.1, then 49.4, then 52.6. Turning the dial up now buys something real.
The consequence is not the one people expect. It is not that you should turn everything up. It is that the bottom of the dial has moved so far that the cheapest setting on the new model lands roughly where the old model's most expensive setting did, at about a quarter of the cost per task. Those two scores sit inside the test's own margin of error, so treat it as a tie on quality rather than a win, and a genuine four-fold cut on price.
Find your floor in twenty minutes
Vendor benchmarks measure someone else's tasks. This measures yours. Do it once and you will know which setting to leave things on.
Pick one job you already know the answer to
2 minutesChoose a piece of work from the last fortnight where you can tell good from bad without asking anyone: a document you summarised, a draft you edited, a spreadsheet you checked. Knowing the right answer is what makes the test readable.
Why this matters
People usually test a new model on something novel and impressive, which tells you nothing, because you cannot grade the output. A job you have already done by hand is the only kind where you can see immediately whether the answer is worse.
Run it once on the highest setting
5 minutesGive it the job with the effort turned all the way up and keep the result. This is your reference, not your target. You are establishing the best it can do so you have something to compare against.
Run the identical job one level down
5 minutesSame wording, same attachments, nothing changed except the setting. Any difference in how you phrase the request will contaminate the comparison and you will not know which change caused what.
Note. Start a fresh conversation rather than continuing the first one. Carrying the earlier answer into the second run lets it copy itself, and both results will look identical for the wrong reason.
Compare the answers before you look at the cost
5 minutesRead both results and decide whether the cheaper one is worse in a way that would matter to the person receiving it. Decide that first, in isolation, because knowing the price gap will bias you.
Why this matters
The honest answer for a lot of business work is that they are indistinguishable. When that happens on your own task, that is far stronger evidence than any published benchmark, because it is your work and you are the one who has to sign it off.
Keep dropping until it breaks, then go back one
5 minutesIf the answer held, go down another level and repeat. The level below the first one that produces a visibly worse result is your floor for that kind of job.
Careful. The floor is per job type, not per business. Summarising a meeting and reviewing a contract for liabilities will not share a floor. Run this again the first time you point it at a genuinely different kind of work.
A starting point for each kind of work
- Summarising and drafting
- Start at the bottom. Meeting notes, first drafts, rewriting something in a different tone, pulling the key points out of a report. If the result reads fine, there is no reason to pay more.
- Checking and comparing
- One or two levels up. Reviewing a document against a set of rules, finding what is missing, comparing two versions. There is a right answer here and you want it found.
- Multi-step work
- Middle to high. Anything where the model has to do one thing, look at the result and decide what to do next. Cutting the thinking short here shows up as it taking the wrong turn early and confidently continuing.
- Long unattended jobs
- High or above. If you are setting it going and walking away, there is nobody there to notice it going wrong. This is the case where paying for more thinking is clearly worth it.
- Anything you will send to a client
- Whatever your test said, plus one level. The cost difference on a single document is pennies and the cost of sending something wrong is not.
Get it to grade its own two answers
If the two results look similar and you cannot tell whether the cheaper one is genuinely as good, paste both into a fresh conversation with this. It is more reliable than reading them side by side, because it forces a specific difference rather than an impression.
Two-run comparison prompt
Below are two answers to the same request, produced by
the same AI at two different effort settings. I do not
know which is which and neither do you.
Compare them on:
1. Factual accuracy. List anything present in one and
missing or wrong in the other.
2. Completeness against the original request.
3. Anything a reader would actually notice.
Do not comment on style or tone unless it changes the
meaning. End with one line: are these materially
different, yes or no, and if yes, name the single
biggest difference.
--- ANSWER A ---
[paste the first result]
--- ANSWER B ---
[paste the second result]Use case one: the long document nobody reads
This is where the cache cut lands hardest, because a long document is re-read constantly while the job runs. It is also the pile of work most businesses have quietly given up on.
Find the document nobody has fully read
Start hereThe supplier agreement, the insurance schedule, the lease, the tender pack, three years of a subscription you have never audited. The test is whether anyone in the business could answer a specific question about page 40 without going to look.
Ask for findings, not a summary
A summary of a long document is close to useless, because it flattens exactly the detail you needed. Ask for something specific instead: every clause that lets the other party change the price, every date that triggers an obligation, every figure that contradicts another figure.
Why this matters
The difference matters more than it sounds. Summarise reduces the document to its most typical content, which is the opposite of what you want. You are looking for the unusual clause, and unusual is precisely what a summary discards.
Make it quote the source for every finding
Require each point to carry the exact wording it came from. This turns a list you have to trust into a list you can spot-check in a minute, and it makes anything invented obvious immediately.
Careful. Spot-check at least three findings against the actual document before acting on any of them. A quote that looks plausible and is not in the source is the failure mode to watch for, and it is not rare enough to ignore.
Run the same request across the whole pile
Once the request works on one document, the marginal cost of running it across every contract you hold is now small enough that the question is no longer whether it is worth it.
Why this matters
This is the practical shape of the price cut. It was never that one document became cheap to analyse. It is that analysing forty of them stopped being a project you have to justify.
Use case two: the job you set going and leave
The clearest change in this release. The model is markedly better at long stretches of unsupervised work, and those stretches are exactly where the cache saving compounds.
Pick something with a checkable result
Start hereReconciling two exports against each other, working through a backlog of records and flagging the broken ones, testing whether a process actually does what the documentation claims. The common feature is that you can verify the output without redoing the work.
Say what done looks like before you start it
Write the finish condition into the request: the file it should produce, the format, what to do when it hits something it cannot resolve. An unattended job with no defined end will keep going or stop somewhere arbitrary.
Why this matters
Anthropic cite one customer, Ramp, running a single unattended job for 38 hours. It found a mistake in their own earlier results, corrected it, and started six further experiments overnight. That only works because the finish condition was clear enough for it to keep deciding what to do next without anyone watching.
Tell it to stop and ask rather than guess
Add an explicit instruction that if it is genuinely unsure, it should stop and record the question rather than pick an answer and continue. Anthropic say this version is better at signalling when it is stuck, which is only useful if you have told it that stopping is allowed.
Note. Without this, a long unattended run tends to make one early assumption and build hours of work on top of it. The failure is not that it gets something wrong, it is that it does not tell you.
Check the first run yourself, in full
Once onlyRead the whole of the first result rather than sampling it. You are not checking this particular output, you are learning where this kind of job goes wrong so you know what to spot-check on every run after it.
The price cut may not reach your bill at all
These prices are what you pay if your business is billed per token, which usually means something a developer has connected up for you. If you use Claude through a monthly subscription instead, you are not paying per token, so the cut does not appear on your invoice. It still matters, because the same work now uses far fewer tokens and your monthly allowance stretches further. Anyone quoting you a per-run cost for an automation should be able to say which of the two applies to you, and if they cannot, that is worth pressing on.
Before you move real work across
- You have run one job you already knew the answer to, at two different effort settings.
- You compared the two answers before you looked at what they cost.
- You know which setting is your floor for at least one kind of work you do often.
- You have spot-checked at least three quoted findings against the original document.
- Any job you leave running unattended has a written finish condition.
- You have told it to stop and ask rather than guess when it is unsure.
- You know whether your business is billed per token or on a monthly plan.
- You have not moved anything client-facing across on the strength of one good result.
Files in this guide
Yours to keep and edit. No attribution needed.