Skip to main content
Modelos de IA

Claude Opus 5.5 Leak: Wait, Test, or Switch Now?

he Claude Opus 5.5 leak claims $4/$20 pricing and a Tuesday launch. None of it is confirmed. Here's what is real this week and how I'd budget, test, or switch.

24 min
Tiempo de lectura
4,669
Palabras
Publicado
Engr Mejba Ahmed

Escrito por

Engr Mejba Ahmed

Compartir Artículo

Claude Opus 5.5 Leak: Wait, Test, or Switch Now?

The Claude Opus 5.5 leak says Anthropic is privately testing a model under the codename claude-wafer-eap, priced at $4 per million input tokens and $20 per million output, with a launch as early as Tuesday, September 22, 2026. As of Monday, September 21, Anthropic has published none of it: no model card, no API identifier, no price sheet. Treat every number as rumor. The things you can actually use this week are StepFun's Step 5 Preview (a live API at $1/$2.70) and Alibaba's Qwen-Image-2.1 (open weights, with a license catch). My advice: don't move Claude budgets on the leak, but do spend an afternoon testing Step 5 Preview.

That's the short answer. The rest of this post covers why I landed there, the math behind it, and the one security trap I found while checking the Step 5 weights.

Over the weekend I watched a video summary go around that packed five labs into eight minutes: an Opus 5.5 leak, a 600B StepFun model, a 3-trillion-parameter MiniMax model, a Qwen 4 reveal, and a Kimi teaser. The auto-transcript spelled half the company names wrong. Anthropic came out as "Enthropic." Qwen came out as "Quinn." That's funny, but it's also a useful warning. When the names are wrong in the summary, the numbers usually got worse along the way too.

So I did the boring part and chased every claim back to its source. Some held up. Several turned out to be one anonymous post repeated by forty sites. One "confirmed" claim is really a CEO's forecast on an earnings call.


The Verification Ledger: Every Claim, Sourced or Flagged

Claude Opus 5.5 leak verification ledger separating rumored claims from confirmed AI model announcements.

This table is what the rest of the post is built on. If you only read one section, read this one.

Claim from the video Status as of Sept 21, 2026 Where it traces
Opus 5.5 in testing as claude-wafer-eap Rumor X posts relayed by aggregators; no Anthropic statement
Opus 5.5 at $4 input / $20 output per 1M Rumor Leak coverage; 20% below Opus 5's official $5/$25
Opus 5.5 has a smaller context window Rumor Some leak write-ups say 872K vs Opus 5's official 1M
Opus 5.5 performs near GPT-6 Astra Rumor Leak claim; no benchmarks exist
Opus 5.5 launches Tuesday, Sept 22 Rumor Same leak chain
Step 5 Preview: 600B total / 27B active, 1M context Confirmed StepFun announcement and docs, Sept 20
Step 5 Preview built for software engineering and finance Confirmed (vendor claim) StepFun announcement
MiniMax M3 Pro at ~3T parameters Company forecast CEO Yan Junjie, interim results call, Aug 26
MiniMax M3.1 focused on reliability and efficiency Unconfirmed No official page, model card, or API ID
Qwen 4 family shown at Apsara Conference Rumor Apsara runs Sept 22–24 in Hangzhou; the agenda doesn't name Qwen 4
Qwen 4 audio models (ASR, TTS, live translation) Unconfirmed Nothing official ties them to Apsara
Qwen-Image-2.1: 7B, open weights, native RGBA Confirmed QwenLM GitHub and Hugging Face, Sept 20
Qwen-Image-2.1 rivals Nano Banana 2 Vendor benchmark only Qwen's own Qwen-Image-Bench
Kimi K3.1 is imminent Rumor Leaks circulating since July 27; no Moonshot confirmation

Three confirmed, one company forecast, one vendor-only benchmark, and nine rumors. That's typical for a week like this. The confirmed stuff is where you can act. The rumors are where you set alerts.

Let's start with the rumor that has your budget attached to it.


What the Claude Opus 5.5 Leak Actually Says

Here's the claim as it circulates. Anthropic is supposedly running a private early-access model called claude-wafer-eap. The "EAP" suffix is read as early-access program. The model was reportedly meant to ship as Opus 5.2 and got renamed 5.5 because the changes were bigger than planned. According to the OrcaRouter write-up tracing the leak, the chain starts with an X account relaying a source called "Lyra," a handle with no track record anyone can verify.

The pricing that got everyone's attention: $4 per million input tokens and $20 per million output tokens, with cache reads reportedly at $0.20 and cache writes at $5. That's exactly 20% below what Anthropic charges for Opus 5 today. WinCentral's coverage lays out the comparison and states plainly that none of it is confirmed.

Then the context window. Some leak write-ups say Opus 5.5 would drop to about 872K tokens, down from the 1M that Opus 5 officially supports. The video framed that as trading length for quality and cost. That's a guess about motive stacked on an unconfirmed number.

What's actually on record:

  • Claude Opus 5 launched July 24, 2026, at $5 input / $25 output per million tokens, with a 1M-token context window and up to 128K output tokens (Anthropic's Opus page).
  • GPT-6 Astra went to approved users on September 3 and reached general availability the next day, at $10 input / $50 output per million on the standard API tier (OpenAI).
  • Neither company has published anything that names Opus 5.5.

Pay attention to what that means for the "near GPT-6 Astra performance" claim. You'd be comparing a model with no published benchmarks against a model that has been out for 17 days. That isn't a comparison. It's a vibe.

Is Claude Opus 5.5 confirmed?

No. As of September 21, 2026, Claude Opus 5.5 is not confirmed. Anthropic hasn't published a model card, API model ID, pricing page, or benchmark for it. The codename claude-wafer-eap, the $4/$20 pricing, the 872K context figure, and the Tuesday launch all come from social posts and the sites that repeated them. The one thing that would confirm it is an Anthropic-issued model ID that returns completions.

I covered a similar leak cycle in my Claude Mythos leak breakdown. Leaks aren't useless. They're a signal about direction, and a bad basis for procurement decisions.

Here's the part I think is actually worth your attention. Suppose the price is real. Anthropic cutting its flagship price 20% within two months of launch would say something about the pressure it's under. That pressure isn't coming only from OpenAI. It's also coming from sparse models that charge a fifth of Opus 5's input price and about a ninth of its output price. Which brings us to StepFun.


Step 5 Preview: The Confirmed Model in This Batch

Step 5 Preview AI model showing confirmed coding, finance, long-context capabilities and lower-cost inference compared with frontier models.

StepFun announced Step 5 Preview on September 20, 2026, and opened API access the same day. The headline specs match the video and come from StepFun directly:

  • 600B total parameters, 27B active per token. Sparse mixture-of-experts, about 4.5% activation.
  • 1M-token context, 64K max output.
  • Input: text, images (up to 60 per request), and video. Output is text only.
  • 92 Transformer layers in what StepFun calls a "narrow-deep" layout.
  • Three reasoning-effort levels (low, medium, high), plus tool calling, JSON Schema output, and prompt caching.
  • Model ID: step-5-preview, served through an OpenAI-compatible chat completions endpoint.
  • Pricing: $1.00 per million input tokens on a cache miss, $0.05 on a cache hit, $2.70 per million output tokens including reasoning.
  • Open weights promised for October 15, 2026. No license named yet.

StepFun positions it for software engineering and professional knowledge work, "with particular strength in finance." The benchmark table (compiled by CellCog from StepFun's release) supports that framing, and it also shows where the model falls short:

Benchmark Step 5 Preview (High) GPT-6 Astra Claude Opus 5 Kimi K3
DeepSWE v1.1 67.7% 74.1% 74.0% 67.5%
Terminal-Bench v4 33.3% 57.9% 52.3% 12.6%
GPQA Diamond 93.5% 96.1% 93.2% 93.5%
HLE 46.5% 54.7% 54.9% 46.9%
FrontierFinance 66.4% 55.0% 69.7% 62.6%

Three things jump out. Step 5 Preview beats GPT-6 Astra on FrontierFinance by more than 11 points while costing a tenth as much per output token. It's roughly tied with Kimi K3 on coding (DeepSWE). And it falls apart on Terminal-Bench, at 33.3% against Opus 5's 52.3%.

That last number matters most to me. Terminal-Bench measures whether a model can drive a shell through a multi-step task without drifting. My daily work is agentic coding in Claude Code, so that's the benchmark closest to what I actually do. A 19-point gap there means Step 5 Preview won't replace Opus in my agent loop. It could replace it in document-heavy analysis, where I'm paying Opus prices to read 400-page PDFs.

Two caveats before you take the table at face value. First, StepFun ran Step 5 in High mode against competitors' Max modes, which works against Step 5, so the real gaps may be smaller. Second, every number comes from StepFun. Artificial Analysis gives the model 44 on its Intelligence Index, level with Kimi K3's Max tier. That's the closest thing to third-party confirmation so far.

The Hugging Face Trap Nobody Is Mentioning

This is the one thing I found this week that I haven't seen anywhere else.

The official weights aren't out until October 15, and StepFun's own Hugging Face org has no Step 5 repo as of today. But search Hugging Face and you'll find TypeSafeAI/Step-5-Preview-BF16, uploaded September 20. It's listed at 604,339.5M parameters, with architecture step3p5v, and it carries the custom_code tag.

That tag is the problem. custom_code means loading the model needs trust_remote_code=True, which runs Python from that repo on your machine. A third-party upload of a model whose weights haven't been released, which asks you to execute its code, is exactly how you'd deliver a payload to someone eager to self-host. I'm not saying this particular repo is malicious. I haven't audited it and won't pretend I have. But it isn't StepFun's, and "wait for the official org" costs you 24 days. The alternative risk is a compromised GPU box.

If you want to try Step 5 now, use the API. That's also where the sparse-model economics show up most clearly, and that's worth understanding properly.


How Sparse MoE Changes What You Pay (and What It Doesn't)

Sparse mixture-of-experts (MoE) is an architecture where a model's parameters are split into many "expert" sub-networks, and a router sends each token through only a few of them. A 600B model with 27B active does roughly the arithmetic of a 27B model on every token, while drawing on the knowledge stored across all 600B.

The rough rule: generating one token costs about 2 FLOPs per active parameter. I ran that across the sparse models I've covered on this blog:

Model Total params Active per token Active share ~GFLOPs per token BF16 weight size
Step 5 Preview 600B 27B 4.5% ~54 ~1.2 TB
MiniMax M3 428B 23B 5.4% ~46 ~0.86 TB
DeepSeek V4.1 Flash (decode) 552B 16B 2.9% ~32 ~1.1 TB
Kimi K3 2.8T 104B 3.7% ~208 ~5.6 TB

A dense 600B model would need about 1,200 GFLOPs per token. Step 5 needs about 54, which is 22 times less compute. That gap is why StepFun can charge $2.70 per million output tokens and Opus 5 charges $25. Anthropic doesn't publish Opus architecture details, so I can't do the same math on Claude. The price gap tells you roughly how far apart the serving costs are.

Here's the catch the video skipped.

Sparse activation cuts compute, not memory. Every one of those 600B parameters has to sit in GPU memory, because the router might pick any expert for the next token. That's 1.2 TB in BF16, or about 600 GB at FP8. You pay for compute per token, but you pay for memory whether or not a token ever arrives.

That changes how the savings reach you, depending on how you run the model:

  1. Through an API, you get almost all of the savings. The provider batches thousands of users onto the same resident weights, so the memory cost is spread across everyone. That's where the $1/$2.70 comes from.
  2. Self-hosted at high utilization, you get most of it. If your GPUs stay busy, 27B-active throughput on 600B-class quality is a great deal.
  3. Self-hosted at low utilization, you may get nothing. A cluster big enough to hold 1.2 TB, sitting idle 80% of the day, can cost more per useful token than just calling Opus.

I went deeper on the prefill/decode split in my DeepSeek V4.1 Flash sparse-activation breakdown. The short version for this post: when the October 15 weights arrive, don't treat "open weights" as "cheap." For most solo developers and small teams, the cheapest way to run Step 5 will still be StepFun's API.

So what does all this mean for an actual invoice?


The Session Math: Opus 5 vs Rumored Opus 5.5 vs Astra vs Step 5

I priced one realistic agent session: 2 million input tokens with 80% served from cache, plus 150,000 output tokens. That's roughly a long refactoring run in which the agent re-reads the same repo context many times. I left out cache-write charges to keep the comparison simple, so real bills will be slightly higher for everyone.

Model Status Per session (cached) Per session (no cache) 100 sessions/day × 22 workdays
GPT-6 Astra Official, prompts under 272K $13.10 $27.50 $28,820
Claude Opus 5 Official $6.55 $13.75 $14,410
Claude Opus 5.5 Rumored pricing $4.92 $11.00 $10,824
Step 5 Preview Official $0.89 $2.41 $1,947

The formula, so you can plug in your own numbers:

# Per-session cost estimate. Prices are USD per 1M tokens.
# Opus 5.5 prices are RUMORED. Swap in official numbers if they ship.
def session_cost(inp, cache_hit_rate, out, price_in, price_cache, price_out):
    uncached = inp * (1 - cache_hit_rate) / 1e6 * price_in
    cached = inp * cache_hit_rate / 1e6 * price_cache
    output = out / 1e6 * price_out
    return uncached + cached + output

PRICES = {
    "opus-5":            (5.00, 0.50, 25.00),   # Anthropic, official
    "opus-5.5-rumored":  (4.00, 0.20, 20.00),   # leak only; unverified
    "gpt-6-astra":       (10.00, 1.00, 50.00),  # OpenAI, <272K prompt tier
    "step-5-preview":    (1.00, 0.05, 2.70),    # StepFun, official
}

for name, (pi, pc, po) in PRICES.items():
    print(name, round(session_cost(2_000_000, 0.8, 150_000, pi, pc, po), 2))

Two things come out of this table that I didn't expect.

First, the rumored Opus 5.5 discount is mostly a caching story. Going from $5 to $4 on input and $25 to $20 on output saves 20%. But the rumored cache-read price falls from $0.50 to $0.20, a 60% cut. On a workload with an 80% cache-hit rate, that brings the per-session saving to about 25%, not 20%. If you run long agent loops over stable context, as most Claude Code users do, a real Opus 5.5 at those prices would help you more than the headline number suggests.

Second, no Opus price cut closes the gap to Step 5. Even the rumored Opus 5.5 costs about 5.5 times as much as Step 5 Preview per session. So the choice isn't "wait for cheaper Opus or switch." It's about which of your tasks actually need Opus-level Terminal-Bench performance. Those tasks stay on Claude whatever the price. Everything else is a routing decision.

That brings me to the image model in this batch, which comes with its own "cheap but read the fine print" story.


Qwen-Image-2.1: RGBA Is the Real Story, the License Is the Catch

Qwen-Image-2.1 native RGBA image generation showing transparent-background output composited into a finished visual design.

Alibaba released Qwen-Image-2.1 on September 20, 2026. I checked the Hugging Face model card directly: 7,115.1M parameters, served through a new QwenImage21Pipeline in diffusers, with 6.5K downloads and 1,228 likes in its first day. The confirmed specs:

  • 7B single-stream diffusion transformer with 32 DiT layers. The same weights handle both generation and editing.
  • Native RGBA output. It writes a real alpha channel through a 64-channel RGBA autoencoder, so there's no separate background-removal step.
  • Up to 10 reference images for multi-subject editing and composition.
  • Native 2K generation, 2048×2048 at 1:1, with other presets such as 2752×1536 at 16:9.
  • Qwen3-VL 8B as the text and condition encoder.
  • Day-0 ComfyUI support with official workflow templates, per the Comfy team.

The RGBA part is the one I care about. For UI and web work, much of image generation's day-to-day pain has been what happens afterward: generate a logo or product shot, run it through a matting model, clean up the fringe, export. Native alpha removes that whole step. If you build landing pages, game sprites, or icon sets, that's a real workflow change.

Now the Nano Banana 2 claim. The video said Qwen-Image-2.1 "rivals" Google's model. Technically, on Qwen's own benchmark (Qwen-Image-Bench), it edges past Nano Banana 2.0 by 0.46 points, about 0.8%. On the same vendor chart, it ranks seventh of 29, behind six closed models led by GPT Image 2.5 Sunburst at 67.01. When the lab that built the model also wrote the benchmark, and the lead is under 1%, that's a tie at best. I'd wait for the independent image arenas to include it before calling it a win.

Then there's the license, which settles it for most readers of this blog. Qwen-Image-2.1 ships under the Qwen Research License Agreement, which bars commercial use without a separate grant. Hugging Face lists it as license:other. That means:

  • For personal projects, research, and internal prototypes: go for it.
  • For client work, SaaS features, or anything that makes money: you need to apply to Alibaba for a commercial grant, and they haven't published pricing.

Earlier Qwen image releases trained people to assume "Qwen means permissive." Don't assume that here. If you're building a paid product, a hosted model with clear commercial terms is still the safer default, even though Qwen-Image-2.1 is the more interesting model.

That leaves the rumor tier: three labs whose next models you can't use yet but that you'll hear about all week.


MiniMax, Kimi, and Qwen 4: Sorting the Rumor Tier

These are the claims to watch, not act on.

MiniMax M3 Pro and M3.1

What's on record: on MiniMax's August 26 interim results call, founder and CEO Yan Junjie said M3 Pro is "expected to scale to approximately 3T" parameters, with more investment in reinforcement learning and long-horizon task training. He also said MiniMax is adapting M3 to domestic Chinese chips and that large domestic compute clusters will start handling production traffic soon. The Information had reported a 2.7T figure in July. "Approximately 3T" doesn't contradict that. It suggests the plan is still moving.

What isn't on record: anything about "M3.1." There's no official page, model card, API ID, or release post. The reliability, inference-efficiency, and agent-generalization themes in the video match what Yan said about company direction, but nobody at MiniMax has attached them to a product called M3.1. Also missing: M3 Pro's active-parameter count, context window, price, license, and release date.

The shipped model is still MiniMax M3: 428B total, 23B active, 1M context, released June 1. My MiniMax M3 first look covers what it does well today.

Kimi K3.1

Moonshot released Kimi K3 on July 16, 2026: 2.8T total parameters, 104B active (16 of 896 experts per token), and 1M context. I covered it in my Kimi K3 review. K3.1 talk goes back to July 27, when a leak claimed an August launch with faster inference, better token efficiency on long reasoning, and stronger coding. August came and went. I couldn't find a primary source for the "cryptic teasers" the video mentioned. The only documented Moonshot teaser was the wordless "3" video before K3 in July. K3.1: unconfirmed, and already a month past its rumored date.

Qwen 4 at the Apsara Conference

The event name checks out. "Aspera conference" in the transcript is Alibaba Cloud's Apsara Conference, which runs September 22–24, 2026, at the Hangzhou International Expo Center under the theme "Intelligence Goes Beyond" (Alibaba Cloud). The history makes the rumor plausible: Qwen 2.5 debuted at Apsara 2024, and Qwen3-Max, Qwen3-VL, and Qwen3-Omni were announced at Apsara 2025.

But the official agenda doesn't name Qwen 4. The closest confirmed signal is Qwen3.8-Flash-Next, released August 26 and described as an early preview of the architecture planned for Qwen 4. Support for that architecture is already landing in llama.cpp, vLLM, and Transformers. The audio lineup the video listed (ASR, TTS, real-time multilingual translation, omni models) fits Alibaba's pattern, but nothing official ties it to this week. I tracked the earlier stealth-model rumors in my Chinese AI models 2026 fact-vs-rumor piece. This week, Apsara either confirms them or it doesn't.

We'll know in about 72 hours. What you do in the meantime is the useful part.


My Playbook for the Next 30 Days: Wait, Test, Budget, Switch

For each of the four moves, here's what I'd do, why, and what could go wrong.

1. Wait: don't touch Claude contracts or routing because of the Claude Opus 5.5 leak

What: Keep your current Opus 5 setup. Don't pre-commit to Opus 5.5 pricing in budgets, client quotes, or annual plans.

Why: Nothing has shipped. Even if a model launches Tuesday, the final price can differ from the leak. So can the context window, and 872K vs 1M matters a lot if you send whole repos. So can the model ID.

What could go wrong: If the model does launch, you'll be a few days behind the early adopters. That's a cheap price for not rebuilding a quote on numbers that turned out to be wrong.

Set this up instead: a one-line check you can run Tuesday morning.

# Lists the model IDs your Anthropic key can see. Nothing new here means nothing shipped.
curl -s https://api.anthropic.com/v1/models \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" | jq -r '.data[].id'

If a new Opus ID shows up, then read the model card, check the context window, and rerun your cost math with the real prices.

2. Test: spend one afternoon on Step 5 Preview through the API

What: Run Step 5 Preview against 10–20 real tasks from your own backlog. Use tasks you've already done with Opus, so you know what a good answer looks like.

Why: At $1/$2.70, testing is almost free. The FrontierFinance and GPQA numbers suggest it could handle analysis-heavy work. The Terminal-Bench number suggests it will struggle with multi-step shell work. Your tasks will tell you where your line falls.

How: The endpoint is OpenAI-compatible, so the standard OpenAI SDK works if you change the base URL. Take the exact base URL from StepFun's platform docs, not from a blog post, this one included.

import os
from openai import OpenAI

# Base URL and key come from the StepFun developer platform (api.stepfun.ai).
client = OpenAI(
    api_key=os.environ["STEPFUN_API_KEY"],
    base_url=os.environ["STEPFUN_BASE_URL"],
)

resp = client.chat.completions.create(
    model="step-5-preview",
    messages=[
        {"role": "system", "content": "You are a senior analyst. Cite the section you rely on."},
        {"role": "user", "content": open("task_07_contract_review.md").read()},
    ],
    # StepFun documents low/medium/high effort levels. Check the docs for the exact parameter name.
)
print(resp.choices[0].message.content)
print(resp.usage)  # log token counts so your cost math uses real numbers

What could go wrong:

  • Data residency. This is a Chinese-hosted API. For client code, contracts, or anything regulated, check your data-processing obligations before sending real data. Use synthetic or public tasks for the first round.
  • Mode confusion. If the effort parameter isn't set correctly, you might be testing low effort against your memory of Opus at max. Log the setting with every run.
  • Unofficial weights. See the Hugging Face warning above. The API is the safe way to test until October 15.

Pro tip: Grade blind. Have a teammate shuffle Opus and Step 5 outputs for the same task and pick the better one without knowing which is which. My early-impression bias has fooled me before, and blind grading catches it.

3. Budget: model the rumored price as a scenario, not a line item

What: Put three columns in your cost sheet: Opus 5 official ($5/$25), Opus 5.5 rumored ($4/$20), and a routed mix, for example 60% of tokens on Opus and 40% on Step 5 Preview.

Why: The routed mix is usually the biggest lever. Using the session math above, moving 40% of sessions from Opus 5 to Step 5 Preview cuts that monthly figure from $14,410 to roughly $9,425. That beats the rumored Opus 5.5 discount on its own, and it doesn't depend on a leak coming true.

What could go wrong: Routing adds real engineering cost. You need a classifier or rules deciding which tasks go where, plus evals to catch quality regressions. If your team can't maintain that, a single provider's price cut may be worth more than a mix you can't keep running. I compared the single-provider approach in my Fable 5.1 price-cut breakdown. The same logic applies here.

4. Switch: only where your own tests say so

What: After the test afternoon, move whole task categories, not individual prompts. "Contract summaries go to Step 5" is a policy you can keep up. "This one prompt goes to Step 5" is a mess.

Why: Category routing is easy to reason about, audit, and undo.

What stays on Claude for me: agentic coding, long shell sessions, and anything where a Terminal-Bench-style failure means a broken deploy. That 19-point gap is too big to route around on price. For the benchmark detail behind that choice, see my Claude Opus 5 benchmarks breakdown, including the five categories where Opus 5 loses.

If you've followed this far, you have a cost model that doesn't depend on any leak being true, which is more than most teams have this week.

If you'd rather have someone build this kind of multi-model routing layer for you, with evals, cost logging, and fallback, I take on exactly that kind of engagement. You can see what I've built on Fiverr.


Where I Could Be Wrong

A few honest limits on everything above.

I haven't run Step 5 Preview yet. The API opened on Sunday. Everything I've said about its strengths and weaknesses comes from StepFun's own benchmark table and one independent index score. My plan is to run it against my own backlog this week. Until then, read my Step 5 section as analysis, not a review.

The leak might be right. Leaks about Anthropic models have been right before. If Opus 5.5 ships Tuesday at $4/$20, my "wait" advice will have cost you roughly 48 hours. I'm fine with that trade. Being early on a real launch earns you very little. Budgeting against a fake one costs you a client conversation.

The session model is simplified. I left out cache-write fees, batch discounts, and GPT-6 Astra's surcharge for prompts over 272K. For your real workload, pull a week of token logs and use the formula above with your actual numbers.

Benchmarks don't measure your codebase. A 67.7% on DeepSWE says nothing about how Step 5 handles your Laravel monolith with its twelve-year-old helpers file. That's why the test step exists.

The unpopular opinion: I think the Opus 5.5 rumor matters less than the Step 5 launch, even though it got ten times the attention. A 20% price cut on one frontier model is incremental. A $2.70-per-million-output model that beats GPT-6 Astra on a finance benchmark is structural. It sets the price that every frontier lab's mid-tier offering now has to justify itself against.


Back to That Misspelled Video

The transcript that started this spelled Anthropic as "Enthropic" and Qwen as "Quinn." Fixing the names took thirty seconds. Checking the claims took far longer. Out of fourteen claims, three were fully confirmed, one was a CEO's forecast, one rested on a benchmark the vendor wrote itself, and nine were rumors, including every number attached to the model with the biggest budget impact.

That ratio is the actual lesson, more than any of the models. In weeks like this, the confirmed releases are the quiet ones and the loud ones are rumors. So here's the challenge for the next 24 hours: pick five tasks you ran through Claude last week, send them to Step 5 Preview, and log the token counts. By Tuesday morning you'll have your own data, and you'll know what an Opus 5.5 launch is worth to you, whether it happens or not.


FAQ

Frequently Asked Questions

Everything you need to know about this topic

claude-wafer-eap is the codename that leaked social posts attach to a rumored Claude Opus 5.5 model in private early-access testing. Anthropic hasn't confirmed the codename, the model, or its specs as of September 21, 2026. For the full sourcing chain, see "What the Claude Opus 5.5 Leak Actually Says" above.

Claude Opus 5.5 pricing is unconfirmed. Leaks claim $4 per million input tokens and $20 per million output tokens, 20% below Opus 5's official $5/$25, with cache reads rumored at $0.20. Anthropic has published no pricing. The session math section shows how the rumored price compares with Opus 5, GPT-6 Astra, and Step 5 Preview.

StepFun Step 5 Preview is available now through StepFun's API, with open weights promised for October 15, 2026. No license has been named yet. Avoid unofficial Hugging Face uploads that need trust_remote_code. Use the API until the official weights arrive.

No, not by default. Qwen-Image-2.1 ships under the Qwen Research License Agreement, which bars commercial use unless Alibaba grants a separate commercial license. Personal, research, and internal prototyping use is fine. For paid products, apply for a grant or use a hosted model with clear commercial terms.

Mixture-of-experts models are cheaper per token because only a small share of parameters runs on each token. Step 5 Preview uses 27B of its 600B, about 4.5%. That cuts compute roughly 22× compared with a dense model of the same size. All the weights still have to sit in memory, though, so the savings show up mostly on high-utilization APIs.

Let's Work Together

Looking to build AI systems, automate workflows, or scale your tech infrastructure? I'd love to help.

Publicidad
Coffee cup

¿Te gustó este artículo?

Tu apoyo me ayuda a crear más contenido técnico detallado, herramientas de código abierto y recursos gratuitos para la comunidad de desarrolladores.

Temas Relacionados

Engr Mejba Ahmed

Engr Mejba Ahmed

Engr. Mejba Ahmed builds AI-powered applications and secure cloud systems for businesses worldwide. With 8+ years shipping production software in Laravel, Python, and AWS, he's helped companies automate workflows, reduce infrastructure costs, and scale without security headaches. He writes about practical AI integration, cloud architecture, and developer productivity.

Artículos Relacionados

Ver Todos

Comments

Leave a Comment

Comments are moderated before appearing.

Learning Resources

Expand Your Knowledge

Accelerate your growth with structured courses, verified certificates, interactive flashcards, and production-ready AI agent skills.

Sample Certificate of Completion

Sample certificate — complete any course to earn yours

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support