Skip to main content
AI Models

GPT 5.5 vs Opus 4.7: I Tested Both. Here's What Won.

Four identical builds on GPT-5.5 and Opus 4.7: 70K vs 250K output tokens, half the wall-clock time, and where Claude's 1M context still wins.

8 min
Read time
1,449
Words
Published
Last revised
Engr Mejba Ahmed

Written by

Engr Mejba Ahmed

Share Article

GPT 5.5 vs Opus 4.7: I Tested Both. Here's What Won.

The tab said 250,000 output tokens. I refreshed the dashboard twice; same number. Opus 4.7 had just finished a solar-system simulation, and GPT-5.5 had finished the identical prompt using 70,000 output tokens — three and a half times fewer — for output that looked, on the surface, nearly equivalent. That gap is why I'm convinced this comparison won't be settled by benchmark sheets. It gets settled by token bills and wall clocks, and the answer depends on what you're shipping. Opus 4.7 is my daily driver, so nobody should mistake what follows for OpenAI cheerleading. The numbers just are what they are.

When GPT-5.5 — "Spud" in the leaks — landed on April 23, roughly six weeks after GPT-5.4, it came with a doubled price ($5 per million input tokens, $30 per million output, up from $2.50/$15) and a token-efficiency pitch to justify it. On Terminal-Bench 2.0 it posted 82.7 against Opus 4.7's 69.4. I wanted to know whether "more with less" survives contact with real work, so I ran four production-style builds through both models with identical prompts, no follow-ups, and full token tracking.

GPT 5.5 vs Opus 4.7: I Tested Both. Here's What Won. - overview of the setup, build 1: personal brand website

The Setup

Four tasks spanning what I actually build in a normal week: a personal brand website (design-heavy, judgment-driven), a solar-system simulation (physics, animation, interactivity), a 3D space shooter (gameplay feel, real-time rendering), and an ecosystem simulation (emergent behavior — the deliberate one-shot killer). One prompt per task, same prompt to both models, and four numbers logged per run: wall-clock time, input tokens, output tokens, estimated cost at list pricing (Opus 4.7 sits around $5/$25, which makes its output tokens slightly cheaper per million than GPT-5.5's).

Harness note, because it shapes everything: GPT-5.5 ran through Codex, Opus 4.7 through Claude Code. Comparing bare APIs would misrepresent both — these models are designed to live inside their harnesses, and as you'll see, the harness is half the economics.

Build 1: Personal Brand Website

Deliberately vague prompt — hero animation, work history, project showcase, contact form, "production-ready" — exactly where autonomous decomposition should shine or collapse.

GPT-5.5 finished in about four minutes for roughly a dollar: coherent, shippable, unspectacular. Opus 4.7 took about fourteen minutes for roughly five dollars: visibly prettier — better type hierarchy, smoother animation — with two real UI bugs (a stuck hover state, a button overflowing its container on mobile). After a QA pass, the two were about equally shippable.

The finding that reframed the whole week: GPT-5.5's token savings didn't come from doing less. Reading the diffs side by side, both implementations covered the same functional surface. Opus wrote more comments, more self-explanation, more breathing room; GPT-5.5 wrote like a senior engineer told to optimize for signal. Fewer output tokens meant compressed style, not reduced scope. I had assumed efficiency would cost completeness. It didn't.

Build 2: Solar System — Where the Cost Thesis Broke

GPT-5.5 won on speed by maybe thirty seconds on an eight-minute task, and then the ledger got interesting: it consumed more input tokens than Opus, because Codex's tool-calling fans out — more file reads, more searches, more context shuffling. Claude Code retrieved more conservatively. Net result: GPT-5.5 slightly faster and about a dollar more expensive, while Opus 4.7's version had better orbital proportions and a palette a designer would actually approve.

This is the experiment that killed my simple "GPT-5.5 wins on cost" model. Output tokens are half the bill. A harness that's aggressive about retrieval can lose on input what the model saves on output. Total economics are a property of the model-plus-harness system, and anyone quoting per-token prices without harness behavior is doing arithmetic, not analysis. It's the same lesson at a different layer as managing token spend inside Claude Code: the invoice is shaped by everything around the model.

Build 3: The Space Shooter That Changed My Mind

I expected Opus to win this one — game logic, real-time controls, my mental file said "deeper reasoning." Instead, GPT-5.5's game was simply better where games live or die: the controls felt tight, projectiles arced naturally, enemies dodged believably, and I caught myself playing it for ten minutes past the point of testing. Under three dollars, about four minutes. Opus 4.7's version had richer sound design and more configurable code, but shipped with input lag, a mistuned firing cooldown, and an enemy-AI freeze bug. I ran both twice to check my own bias. Same result.

The generalizable insight: GPT-5.5 commits to aggressive defaults — timing curves, response feel, feedback loops — while Opus hedges toward flexibility, adding configuration where a player just needed a good default. For products where defaults are the experience, that hedging is a real cost. No leaderboard would have predicted this, which is rather the point.

Build 4: The Ecosystem Simulation Both Models Failed

Predators, prey, plants, hunger, aging, reproduction, and a population that must reach stable equilibrium unaided. I include this prompt because both models always fail it, and how a model fails is more informative than how it succeeds.

GPT-5.5's failure was math-shaped: structurally correct simulation, predators starving out in thirty seconds because the hunger and reproduction constants were mistuned. Five minutes of parameter fiddling from working. Opus 4.7's failure was structural: a reproduction-trigger bug — every entity above a fitness threshold reproduced every tick — that overran the screen and collapsed the framerate. Fixing it meant refactoring the loop, twenty minutes minimum.

That asymmetry repeated all week and now factors into my routing: when GPT-5.5 misses, the miss tends to be tunable; when Opus misses on a hard one-shot, the miss tends to be architectural. Tunable failures are cheap. Architectural failures cost real time. And both models burned $3-5 producing broken simulations, so budget two or three iteration cycles for genuinely novel work — anyone selling one-shot-to-production is selling the demo.

The Aggregate

Across all four builds: GPT-5.5 finished in roughly 20 minutes 49 seconds of total runtime against Opus 4.7's 40 minutes 43 seconds. About 70K output tokens against 250K. Input tokens slightly higher for GPT-5.5 (roughly 2.7M vs 2.5M). Total cost: GPT-5.5 marginally cheaper — by about three dollars — despite the doubled unit price.

So the headline efficiency claim is real. And the fine print matters: if your workload is input-heavy — big codebase loads, long documents — Codex's retrieval appetite erodes the advantage.

One correction to the discourse, because it's load-bearing: people keep saying "GPT-5.5 is capped at 400K context." Inside Codex, yes — the harness runs a 400K window. But the API ships GPT-5.5 with a 1M-token context, OpenAI's first. The practical comparison for daily work is still harness-level, though, and there Claude Code with Opus 4.7's million-token window lets me do things Codex currently can't — like loading the better part of a large Laravel monolith in one working set. I've written about what a 1M context actually changes in practice; short version: context size is a capability, not a spec-sheet line, and for codebase-wide refactoring it's still Anthropic's moat in my stack.

What I'm Actually Doing With This

My routing after the test week, stated as rules:

GPT-5.5 via Codex for fast prototyping that fits its harness window, for anything where moment-to-moment feel or micro-interaction quality decides the outcome, and for vague prompts on well-trodden problem shapes where its confident defaults save a clarification round.

Opus 4.7 via Claude Code for the large production codebase where the context window is doing irreplaceable work, for design-taste-sensitive builds, and for long agent sessions where retention across a big working set matters more than sprint speed.

Both in parallel when the task is hard enough that two competing solutions are worth the bill. I already run Claude Code and Codex side by side on the same repo, and this test hardened my view that the endgame isn't picking a winner — it's routing sub-tasks to whichever model's failure mode you can afford. My longer Codex vs Claude Code field test covers the harness differences that make that practical.

On the benchmarks: GPT-5.5's 13-point Terminal-Bench lead is a real signal and it genuinely didn't predict my results. It said nothing about Opus's better visual proportions in build two, nothing about GPT-5.5's superior game feel in build three, nothing about the tunable-versus-architectural failure split in build four. Benchmarks measure the harness-free model on benchmark-shaped tasks. You ship neither of those things.

The feeling I didn't expect at the end of the week was relief. Not because one model won — because both are now good enough that choosing between them is workflow optimization rather than quality compromise. The six-week release cadence is exhausting, and my coping strategy is the one I'd hand you: keep three real tasks you've already done, rerun them on each new model, decide from your own numbers, and ignore everything else. Don't trust my four experiments either; the only comparison that binds you is the one run on work you have already shipped. The head-to-head prompts I reuse for these tests live in my prompt library.

Coffee cup

Enjoyed this article?

Your support helps me create more in-depth technical content, open-source tools, and free resources for the developer community.

Related Topics

Engr Mejba Ahmed

Engr Mejba Ahmed

Engr. Mejba Ahmed builds AI-powered applications and secure cloud systems for businesses worldwide. With 8+ years shipping production software in Laravel, Python, and AWS, he's helped companies automate workflows, reduce infrastructure costs, and scale without security headaches. He writes about practical AI integration, cloud architecture, and developer productivity.

Related Articles

Browse All

Comments

Leave a Comment

Comments are moderated before appearing.

Learning Resources

Expand Your Knowledge

Accelerate your growth with structured courses, verified certificates, interactive flashcards, and production-ready AI agent skills.

Sample Certificate of Completion

Sample certificate — complete any course to earn yours

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support