Skip to main content
Claude Code

Claude Token Limits: It's Context Rot, Not a Cap

Why you keep hitting Claude token limits: context rot, not a cap. The hygiene habits that cut my real session costs on a production codebase.

7 min
Read time
1,360
Words
Published
Last revised
Engr Mejba Ahmed

Written by

Engr Mejba Ahmed

Share Article

Claude Token Limits: It's Context Rot, Not a Cap

People who search for "Claude token limits" are usually asking the wrong question. I know because I asked it too. You hit the wall mid-task, so you assume the problem is the ceiling and go hunting for a bigger window or a higher plan. After a year of daily Claude Code on a production Laravel codebase, including weeks where I ran ten-agent workflows against a 581-post content database, my conclusion is different: the limit that actually hurts you arrives long before the cap does, and it is called context rot. Quality degrades as the window fills. The cap just makes the degradation official.

Claude Token Limits: It's Context Rot, Not a Cap - overview of the two limits people conflate, context rot is measurable, not a vibe

The two limits people conflate

There are two separate things called "Claude token limits," and they need different fixes.

The context window is how much a single conversation can hold. When people complain Claude "forgot" something from earlier in a long session, or watch /compact fire automatically, this is the limit they are touching.

Usage limits are plan-level budgets that reset on a schedule. I learned my account's reset times empirically, because two of my batch-rewrite workflows died mid-draft at session limits, and the work could only resume at the next reset (for my account, fixed local times, roughly five hours apart). If you are hitting these, the fix is workload shaping and model choice, not context tricks.

Most frustration attributed to the second limit is caused by mismanaging the first. A bloated context makes every message expensive, which burns your usage budget faster, which makes you hit the plan limit sooner. Fix the hygiene and both limits recede.

Context rot is measurable, not a vibe

The strongest evidence that "more context" is not free came from Chroma's research report, "Context Rot: How Increasing Input Tokens Impacts LLM Performance." They evaluated 18 models, including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, and found that performance does not hold steady as input length grows, even on tasks the models handle easily at short lengths. Models do not process the 150,000th token with the same reliability as the 1,000th.

That matches exactly what I see in practice. The failure is never a clean "out of memory" error. It is a session that starts confidently referencing a file we discussed an hour ago and gets a detail subtly wrong. On a codebase with real conventions (this one has a rule that composer test runs PHPStan, not PHPUnit), subtle wrongness is more expensive than obvious failure, because you do not catch it until the diff review.

So the goal is not "stay under the cap." The goal is: keep the working context small enough that everything in it is still being processed reliably. Everything below follows from that.

The hygiene habits I actually run

These are not theoretical. This is the configuration and muscle memory on the machine I am typing on.

Run /context before the work, not after the wall

/context shows what is consuming your window right now: system prompt, tools, MCP servers, files, conversation. I run it at the start of any session that will be long, because the expensive discovery is always something passive. Tool definitions and MCP servers cost tokens on every single message whether you use them or not.

Carry one MCP server, not nine

My .mcp.json for this project declares exactly one server: Laravel Boost. That is a deliberate endpoint of an evolution that started with many. Every server you connect injects its tool schemas into every message. Boost earns its slot because search-docs, tinker, and database-schema replace whole exploratory conversations. The others got replaced by CLIs, which cost tokens only when invoked. My rule now: an MCP server must save more context than its schemas consume, measured over a normal session, or it gets demoted to a CLI.

/clear aggressively between tasks

Unrelated task, fresh window. The residue of a debugging session adds nothing to a blog-content task except noise and cost. It took me embarrassingly long to internalize this because a long transcript feels like accumulated value. It is not. The durable value belongs in files, which is the entire subject of my six levels of Claude Code memory.

Slice big work into subagents with their own windows

This is the biggest lever and the least used. When I triaged 581 blog posts for indexation problems, no single context could hold the corpus. The structure that worked: ten research agents, each assigned a slice, each with its own full window, each writing findings to report files; then ten rewrite agents consuming those reports. The orchestrator's context stays tiny because it holds pointers, not payloads. The filesystem is the shared memory. I documented the full pattern in my agent swarm architecture post.

The subtle win: subagent slicing is context-rot immunity, not just parallelism. Each worker operates in the early, reliable region of its own window instead of the degraded tail of a shared one.

Hand off through manifests, not through /compact

/compact summarizes the conversation to reclaim space, and it is fine for mid-task breathing room. But for work that must survive, I trust a written manifest over a lossy summary. During my 84-post SEO rewrite, a REWRITE-STATUS.json file tracked every post's state, updated after each unit of work. When sessions died at usage limits, the next session resumed from the manifest in one step. Compaction is memory you hope survived. A manifest is memory you wrote down.

Keep the always-loaded layer lean

CLAUDE.md and anything else loaded automatically is a tax on every message. Mine holds only project-specific gotchas a competent developer would get wrong on day one. Generic advice gets deleted. If you want the aggressive version of this philosophy, the Caveman approach to token optimization pushes it further than I do, but the direction is correct.

What about the giant context windows?

Large windows are genuinely useful for genuinely large single artifacts: a huge log file, a long legal document, a full codebase read. They are not a substitute for hygiene, because of the rot curve. Filling a million-token window and expecting uniform recall is exactly the assumption the Chroma research falsifies. My take on when the big window earns its cost is in managing Claude Code's 1M context, but the summary is: use it for payloads, not for long-running conversational accumulation.

The cost arithmetic nobody does

Here is the compounding effect that makes hygiene financially real. In an agentic session, prior conversation is reprocessed as the session continues, so a fat context does not cost you once. It costs you on every subsequent message. Letting 40,000 tokens of stale exploration sit in the window while you send 30 more messages means paying for a large chunk of that waste dozens of times. This is why two developers on identical plans have wildly different experiences of "Claude token limits": one of them is dragging a full transcript everywhere, the other keeps a working set. When I hear "the limits are unusable," I now translate it to "my sessions have no hygiene," and I say that as someone whose own habits had to be corrected by billing reality. My slash-command routine for this, /context, /clear, and friends, is written up in the slash commands I use daily.

The honest priority order

If you are hitting Claude token limits today, do these in order, and stop when the pain stops:

  1. Run /context and delete passive consumers (unused MCP servers first).
  2. Start /clear-ing between unrelated tasks, today.
  3. Move durable facts out of the transcript into CLAUDE.md or memory files.
  4. Split batch work into subagents with per-worker windows and file-based handoffs.
  5. Track long tasks in a manifest so a dead session costs nothing.
  6. Only then think about plans and bigger windows.

Most people never need past step 4. I run all six because content operations at the scale of hundreds of posts leave no choice, and the system has survived session deaths, usage resets, and multi-day gaps without losing state.

I also do this professionally: auditing AI-assisted development workflows, including token-cost and context architecture for teams whose Claude bills grew faster than their output. If your team's sessions are hitting walls that hygiene should have prevented, my services page explains how I work with teams like yours.

Coffee cup

Enjoyed this article?

Your support helps me create more in-depth technical content, open-source tools, and free resources for the developer community.

Related Topics

Engr Mejba Ahmed

Engr Mejba Ahmed

Engr. Mejba Ahmed builds AI-powered applications and secure cloud systems for businesses worldwide. With 8+ years shipping production software in Laravel, Python, and AWS, he's helped companies automate workflows, reduce infrastructure costs, and scale without security headaches. He writes about practical AI integration, cloud architecture, and developer productivity.

Related Articles

Browse All

Comments

Leave a Comment

Comments are moderated before appearing.

Learning Resources

Expand Your Knowledge

Accelerate your growth with structured courses, verified certificates, interactive flashcards, and production-ready AI agent skills.

Sample Certificate of Completion

Sample certificate — complete any course to earn yours

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support