Skip to main content
Claude Opus 5

Claude Opus 5 vs Fable 5: Half Price, Not Half Cost

Claude Opus 5 vs Fable 5 tested across nine real knowledge-work tasks — per-task cost, clock time, and why half the token price never halved the actual bill.

15 min
Tempo de leitura
2,960
Palavras
Publicado
Engr Mejba Ahmed

Escrito por

Engr Mejba Ahmed

Compartilhar Artigo

Claude Opus 5 vs Fable 5: Half Price, Not Half Cost

Claude Opus 5 vs Fable 5: Half Price, Not Half Cost

Ten and a half hours. Two million output tokens. That's what one model burned working through nine knowledge-work tasks. The other model finished the same nine in three hours and twenty minutes on 832,000 tokens — and it was the one costing twice as much per token.

So here's the short answer for anyone comparing Claude Opus 5 vs Fable 5 on cost: Opus 5 is priced at $5/$25 per million input/output tokens against Fable 5's $10/$50, but it spends so many more tokens getting to an answer that the real-world saving lands closer to 20–26% per task, not 50%. Opus wins decisively on technical accuracy and self-verification. Fable wins on speed, visual polish, and anything with a design surface. Which one is cheaper depends entirely on what you hand it.

That gap between the sticker price and the invoice is the whole story, and it took nine tasks with actual dollar receipts to see it.

Let me be straight about whose receipts those are before we go further.

Where These Numbers Came From

I didn't run all nine tasks myself. The battery — semantic-search diagrams, bug hunts in a live codebase, a viral video edit, a landing page, a LinkedIn carousel, audience research, a slide deck, a computer-use game, and a structural physics simulator — comes from a head-to-head evaluation I've spent the past week pulling apart, task by task and receipt by receipt.

What I did do is check every claim in it against what's independently published, because a single evaluator's run is an anecdote until something corroborates the mechanism. And this one corroborates unusually well, which is why I'm writing it up instead of scrolling past.

The published anchors, all verified as of late July 2026:

Metric Claude Opus 5 Claude Fable 5 Source
Input / output per 1M tokens $5 / $25 $10 / $50 Anthropic
Artificial Analysis Intelligence Index (max effort) 61 60 Artificial Analysis
Average cost per Intelligence Index task $2.03 $2.75 Artificial Analysis
AA-Briefcase (agentic knowledge work, Elo) 1,720 1,574 Artificial Analysis
Frontier-Bench v0.1 (agentic terminal coding) 43.3% 33.7% Anthropic launch table
CursorBench 3.2 70.1 70.4 Cursor
SWE-bench Pro 79.2% 80.0% Anthropic launch table

Look at rows two and three together, because that pairing is the argument in miniature. On raw intelligence the two models are effectively tied — 61 against 60. On cost per task, Opus comes in at $2.03 against $2.75. That's a 26% saving, on a model priced 50% lower per token. Roughly half of the advertised discount evaporates somewhere between the price sheet and the finished task.

Artificial Analysis also clocked Opus 5 at max effort generating 100 million output tokens across its Intelligence Index evaluation, against a cross-model median of 63 million. Their own word for it is "very verbose." Meanwhile the developer Theo, running his own real workloads, put the practical saving at "20 to 25% off" and flagged that Opus is slower than Fable and fills context windows faster.

Three independent sources. One pattern. Now watch it show up in nine specific deliverables.

Claude Opus 5 vs Fable 5 on Code: 4 of 4 Against 2 of 4

The evaluation ran bug identification and codebase exploration twice per model — identical prompts, identical codebase, four planted failure conditions, scored out of 95 with an independent pass from Codex as the judge.

Run Model Time Cost Tests passed Score
1 Opus 5 13 min $4.22 4 / 4 93 / 95
2 Opus 5 ~20 min $6.50 4 / 4 93 / 95
1 Fable 5 11 min $5.30 2 / 4 66 / 95
2 Fable 5 12 min $8.73 2 / 4 66 / 95

Read the "tests passed" column twice. Fable 5 — the flagship, the expensive one, the model positioned as the deep-reasoning ceiling — caught half the bugs Opus caught. Twice. On the same codebase.

And the qualitative note from the evaluation makes it worse for Fable in a specific, recognisable way: Fable's patch was cleaner. Tighter diff, better naming, more pleasant to read in review. It just didn't fix as much. Opus produced the uglier, more thorough patch that actually held under test.

I've reviewed enough AI-generated pull requests to know exactly how that failure mode kills you. A clean diff reads as correct. It sails through review because nothing in it looks wrong. The messy patch that covers the edge case gets three comments asking why it's so defensive — and it's the one you should merge.

This is the row I'd point at if someone asked why Opus 5 is my default for anything touching a repository. It also cost less than Fable on both runs while taking longer, which is a preview of every trade-off that follows.

Hold onto that word "thorough." It's about to become expensive.

Claude Opus 5 vs Fable 5 on Creative Work: Fable Wins on Sight

Flip the task type and the ranking flips with it.

Viral video announcement (hyper-edited). Opus produced two versions — one vertical, one landscape — in 40 minutes for $11.19. Fable produced one, in 7 minutes 49 seconds, for $7.00, and included music and sound effects Opus didn't reach for. Opus's output was described as slightly computerised and less professional. Fable's stylistic choices were better, though it introduced factual inaccuracies traceable to stale context in its window.

YouTube outline plus slide deck on context engineering for AI agents. This one isn't close.

Model Time Cost Output
Fable 5 8 min $40.71 29 slides, visually appealing, on-brand
Opus 5 1 hr 15 min $33.00 More basic deck, lacked branding polish

Opus was cheaper by $7.71 and took nine times longer to produce something the evaluator wouldn't present. That's the cost calculus most comparisons miss entirely: seven dollars saved against sixty-seven minutes lost and a deck you now have to fix by hand. If your time has any value at all, Fable won that task on price too.

LinkedIn post and carousel. Both models landed on the same thematic conclusion — audience distrust of AI agents rising even as capability improves. Opus built a branded carousel for $8.22 and used significantly more tokens. Fable produced a tweet-style carousel in about 7.5 minutes for $6.17, and it was the one the evaluator preferred visually. Cheaper and better on a task with a design surface, from the model with double the per-token price.

Marketing landing page. Roughly 60 minutes and $35.83 for Opus, about 22 minutes and $20.50 for Fable. Both landed high quality and on-brand. Opus was accurate but wordy and generic; Fable shipped nicer animations. Neither definitively won, and both needed manual tweaks — which is the honest outcome for AI-generated landing pages in July 2026, whatever anyone's demo video shows you.

Semantic search diagram (Excalidraw). Fable's was more visual, less organised, with minor inaccuracies. Opus's was more structured and more detailed. For teaching a concept, Opus's version was preferred. For putting on a slide, Fable's.

Five creative tasks, and the pattern holds without a single exception: Fable is faster and prettier, Opus is more accurate and more verbose. Not once did the ranking cross over.

Which raises the question that actually decides your bill.

Does Claude Opus 5 Cost Less Than Fable 5 in Practice?

No — not reliably. Claude Opus 5 costs half as much per token as Fable 5, but it consistently spends two to three times more tokens and time to complete the same task, so the real-world saving lands around 20–26% rather than 50%, and it disappears entirely on creative and visual work.

Here's the consolidated ledger from all nine tasks:

Total across 9 tasks Claude Opus 5 Claude Fable 5
Active working time ~630 min (~10.5 hrs) ~200 min (~3.3 hrs)
Output tokens ~2,000,000 ~832,000
Price per 1M output tokens $25 $50
API calls per session comparable comparable
Tool calls per session comparable comparable

Multiply it out. Two million Opus tokens at $25 costs $50. Eight hundred thirty-two thousand Fable tokens at $50 costs $41.60. The half-price model produced the larger output bill — while also consuming three times the wall-clock time.

The detail I keep coming back to is that API calls and tool calls per session were comparable between the two. Opus wasn't taking more actions. It was writing far more text per action. That distinction narrows the explanation down to one thing, and it's the most useful finding in the entire exercise.

Why Opus 5 Burns Tokens: The Verification Loop

Opus 5's defining behaviour is that it checks its own work. Constantly. Unprompted.

Artificial Analysis documented this directly: Opus 5 verifies its own output without being asked, and prompt instructions like "include a final verification step" or "use a subagent to verify" now trigger over-verification. Their finding is blunt — removing those instructions cuts token consumption with no measured loss in quality.

Read that again if you write prompts for a living. Verification instructions that were free quality insurance on Opus 4.8 are now a tax you're paying twice.

The nine-task data shows both faces of that behaviour in the same afternoon. On the bug hunt, verification is precisely why Opus caught 4 of 4 while Fable caught 2 — it ran the tests, read the failures, and iterated. On the slide deck, the identical instinct produced 75 minutes of self-review on a task where nobody needed a second opinion about slide 19's bullet spacing.

The structural simulator task makes the trade-off unmissable. Both models built a physics simulator that adds nodes and weights and stress-tests structures against weather load:

Version Character Cost Time
Version 1 (Opus) Overwhelming, dense UI, legacy-software look, multiple nodes/beams/loads $112 2 hr 26 min
Version 2 (Fable) Simpler, user-friendly UI, drag-and-drop node placement, more AI-generated look $73 7 min

Two hours and twenty-six minutes against seven. A hundred and twelve dollars against seventy-three. Opus spent that budget on extensive verification and iterative testing loops, and what it bought was a denser, more capable, considerably less pleasant tool. Fable spent seven minutes and shipped the one a human would rather use.

Neither is wrong. They're answers to different questions, and the model can't tell which question you meant.

If wiring up model routing like this across a real production workflow sounds like the kind of unglamorous plumbing you'd rather hand off, it's exactly the work I take on — you can see what I build here. The routing rule itself is simple enough to implement yourself, and it's coming up in a moment.

First, the test that fell apart.

The Snake Game Test That Broke — And Why It's the Most Useful Result

Task eight was computer use: play Google's Snake game through a browser and score well.

The evaluator ran Opus 5 twice by accident. Fable never got its turn. In most write-ups this test quietly disappears — and if you've ever run a benchmark suite at 1am you know exactly how it happens.

I'm glad it survived, because the botched test produced better information than the clean one would have.

Two runs. Same model, same prompt, same game. One run followed the instructions poorly and played far more games than it was asked to. Average scores across the runs ranged from roughly 63 in one to over 2,144 in another.

That's a thirty-four-fold spread on identical inputs.

Every table above — mine, Anthropic's, everyone's — implies a determinism that doesn't exist. When a coding run costs $4.22 once and $6.50 the next time on the same codebase, that's a 54% cost variance nobody controlled for. The scores happened to be stable at 93/95 both times, which is genuinely reassuring about Opus's accuracy. The cost was not stable at all.

So treat every single figure here as one draw from a distribution, not a specification. That includes the ones flattering the model I'm recommending.

The Routing Rule I'd Actually Run

Nine tasks, two models, and the conclusion isn't "pick one."

Send to Opus 5: anything with a correctness condition. Bug hunts, codebase exploration, refactors, data work, security review, migrations, anything where a wrong answer is expensive and a slow answer isn't. The verification loop is the product. Pay for it where it pays you back.

Send to Fable 5: anything with a taste condition. Video edits, slide decks, social carousels, landing pages, interfaces a human has to enjoy using. Fable is three times faster with better stylistic instincts, and on this evidence its per-task cost on creative work often lands below Opus despite double the token price.

Send to neither: the routine content, the summaries, the reformatting, the boilerplate. Both of these models are frontier-priced overkill for tasks a Sonnet 4.5-class model handles at a fraction of the cost. Matching model intelligence to task difficulty is still the largest cost lever available, and I've laid out the full ladder in my AI agent cost optimization guide.

The structure I've settled on runs Opus as an orchestrator rather than a worker. Opus plans and decomposes, delegates the generation to a cheaper or more stylistically suited model — Fable for design-facing subtasks, Sonnet for routine ones — and then does what it's genuinely best at: reviewing the result and catching what's wrong. You get the verification moat without paying Opus to write every intermediate token, and session context stays in the delegate model rather than being rebuilt on every hop. It's the same delegation pattern I sketched out in the Claude Code agent teams playbook, applied across model vendors instead of across subagents.

Two prompt-level adjustments do most of the remaining work:

  1. Strip verification instructions from Opus prompts. "Verify your work," "double-check," "use a subagent to confirm" — these now cause over-verification and burn tokens for nothing measurable. Opus already does this unprompted.
  2. Sweep the effort ladder before you settle. Opus 5's output token usage spans roughly 8x from low to max effort. Low scores 51 on the Intelligence Index, max scores 61. If your task lives in territory where 56 is plenty, medium effort gets you there at a fraction of max's token spend. I walked through the mechanics of this in my Claude Opus 5 benchmarks breakdown, and it's the single highest-impact change after switching models.

Then monitor. Actual token consumption per task type, weekly. The routing rule that's right today shifts with the next model release, and the only way to know is to watch your own numbers — which is the discipline underneath everything in how I cut Fable 5 usage costs.

The Part That Should Make You Uncomfortable

Three honest caveats, because a comparison without them is marketing.

Opus 5 hallucinates more when uncertain. Artificial Analysis measured its hallucination rate rising 14 points to 50% compared to Opus 4.8 — the model now answers more often instead of declining. A model that verifies its own work and also confabulates more confidently under uncertainty is a specific hazard: the verification loop can validate a fabricated premise. On the bug hunt this didn't bite, because tests are ground truth. On research tasks with no external check, it very much can.

The prompt matters more than the model. Across all nine tasks, the evaluator's recurring note was that prompting quality, context freshness, and workflow integration moved outcomes more than the model choice did. Fable's video inaccuracies came from stale context, not weak reasoning. A well-prompted Fable beats a badly-prompted Opus on tasks Opus should own — which makes the benchmark deltas far less decisive than the tables suggest. My Fable 5 prompting habits post covers the specific patterns that close most of that gap.

Nine tasks is not a study. One evaluator, one codebase, one brand context, no repetition on seven of the nine tasks — and the two tasks that were repeated produced a 54% cost swing. Every number here should move your priors. None of them should settle an argument.

What survives all three caveats is the mechanism, not the magnitude. Opus verifies more, so it writes more, so it costs more per task than its token price implies. That relationship is confirmed by Anthropic's own model behaviour, by Artificial Analysis's token counts, by an independent developer's production workloads, and by nine deliverables with dollar amounts attached. The exact percentages will move. The direction won't.

What This Actually Changes

Go back to that opening ratio: 630 minutes against 200. Two million tokens against 832,000. A model priced at half, spending more than double.

Anthropic's headline — near-Fable intelligence at half the price — is true about the price sheet and misleading about your invoice, and the distance between those two things is measured in verification loops. Opus 5 doesn't cost less because it's cheaper. It costs less sometimes, on some tasks, when its thoroughness is buying you something you actually needed.

Here's what I'd do before Monday: take the last ten tasks you gave a frontier model and sort them into two piles — the ones with a correctness condition and the ones with a taste condition. Then check what you spent on each pile. My guess is you'll find you've been paying a verification premium on work nobody was ever going to check, and paying it at frontier rates.

The models aren't competing. Your routing table is just missing.

FAQ

Frequently Asked Questions

Everything you need to know about this topic

Opus 5 is better at technical accuracy, verification, and long-running agentic work; Fable 5 is better at speed, visual design, and creative output. They score 61 and 60 on the Artificial Analysis Intelligence Index — effectively tied overall. The right answer depends on whether your task has a correctness condition or a taste condition.

Opus 5 is 50% cheaper per token ($5/$25 vs $10/$50 per million) but only about 20–26% cheaper per completed task, because it consumes substantially more tokens. Artificial Analysis measured $2.03 per task for Opus 5 against $2.75 for Fable 5.

Opus 5 verifies its own work unprompted, running iterative self-review loops that generate large volumes of output tokens. Artificial Analysis recorded 100M tokens at max effort against a 63M median. Adding explicit verification instructions to prompts now causes over-verification and wastes tokens.

For design-facing front-end work, yes — Fable produces cleaner, more polished interfaces faster. For bug-finding and correctness-critical work, the nine-task evaluation showed Fable catching 2 of 4 planted bugs against Opus's 4 of 4, on identical prompts and codebase.

Fable 5 is substantially faster in wall-clock time. Across the same nine tasks, Fable finished in roughly 200 minutes of active working time against Opus 5's 630 minutes — about three times quicker, with comparable API and tool call counts per session.

Let's Work Together

Looking to build AI systems, automate workflows, or scale your tech infrastructure? I'd love to help.

Publicidade
Coffee cup

Gostou deste artigo?

Seu apoio me ajuda a criar mais conteúdo técnico aprofundado, ferramentas open-source e recursos gratuitos para a comunidade de desenvolvedores.

Tópicos Relacionados

Engr Mejba Ahmed

Engr Mejba Ahmed

Engr. Mejba Ahmed builds AI-powered applications and secure cloud systems for businesses worldwide. With 10+ years shipping production software in Laravel, Python, and AWS, he's helped companies automate workflows, reduce infrastructure costs, and scale without security headaches. He writes about practical AI integration, cloud architecture, and developer productivity.

Discussion

Comments

0

No comments yet

Be the first to share your thoughts

Leave a Comment

Your email won't be published

7  +  11  =  ?

Artigos Relacionados

Ver Todos

Comments

Leave a Comment

Comments are moderated before appearing.

Learning Resources

Expand Your Knowledge

Accelerate your growth with structured courses, verified certificates, interactive flashcards, and production-ready AI agent skills.

Sample Certificate of Completion

Sample certificate — complete any course to earn yours

Engr Mejba Ahmed

Engr Mejba Ahmed

Claude Code Expert · Online

👋

Hey there!

Quick Actions

WhatsApp Instant reply

Chat on WhatsApp

+880 1723 741224 · Instant reply

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

[email protected]

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support