The Real Reason Your Claude Limit Fills Up Fast
To get straight to the point, it's not that the model got dumber; it's that my overhead grew.
But simple surface-level tips like "shorten your CLAUDE.md" aren't enough. You need to understand the structure of why tokens leak to truly stop it.
(Honestly, I know many AI beginners might find this hard to understand.
So, I've included prompts at the end of this post that even beginners can use.
If you don't understand the technical parts, just copy and paste those. I hope you get at least something out of this!)
The Core Mental Model (Understand this, and you've got 90% of it)
Transformers re-process the entire conversation from the beginning every single turn.
When you send your 30th message, the model reads: โ Messages 1โ29 + all answers โ All tool call results (PR diffs, file reads, etc.) โ CLAUDE.md โ System Prompt โ MCP tool definitions โ + your 30th message.
It processes all of this before it even starts answering.
In other words, the 30th turn isn't just 30 times the 1st turn; it's the sum of everything accumulated processed every time.
Starting from here, you can naturally see where the tokens are leaking.
9 Holes Where Tokens Leak
The percentage figures in the original source (14%, 13%...) are from one specific case, so generalization is risky. I've reorganized them by impact level.
1. CLAUDE.md Bloat โ Impact โ โ โ
This is included in every message as long as the session is active. It's not lazy-loaded. A 2,000-token CLAUDE.md processed 200 times for 200 messages = massive waste. Official recommendation: Under 200 lines, 300โ600 tokens.
2. Conversation Accumulation โ Impact โ โ โ
Exactly as the mental model describes. It's not weird that your limit hits 60% after two or three PR reviews; it's structural.
3. Tool Output Accumulation โ Impact โ โ โ
Fetching a PR diff once can inject thousands of lines. Reading 20 files means those 20 files follow you until the end. This is a more accurate culprit than the "hooks" mentioned in other posts.
4. Cache Misses โ Impact โ โ
Prompt caching is applied automatically but expires if not used for a short period. If you frequently edit CLAUDE.md mid-session, the cache breaks every time.
5. Skills โ Impact โ (The original source was slightly wrong here)
Skills are only loaded when called. Only the metadata stays resident. The real problem is when a single skill becomes bloated.
6. "Just in Case" MCPs โ Impact โ โ
If you have 12 MCPs connected, 12 tool definitions are injected into every call. Keep only the 3 you actually use as active.
7. Extended Thinking Default โ Impact โ โ โ
Usually ON by default. The base budget can go up to tens of thousands of tokens (charged as output). Using deep reasoning just to change a variable name is a huge waste.
8. Watching a Wrong Answer to the End โ Impact โ โ
If the answer is going off the rails, stop it immediately. If you don't, that entire output becomes input for the next turn.
9. Cumulative Notifications/Meta Messages โ Impact โ
Small, but they are "quiet offenders" when they accumulate.
Always Diagnose Before Fixing
This is the part people miss.
/context โ Displays tokens by item in the context
/usage โ Session usage
/cost โ Cumulative API cost
Running /context just once will show you the #1 leak in your specific case within 5 seconds.
Most results are similar:
- Accumulated tool outputs (overwhelmingly #1)
- CLAUDE.md
- MCP tool definitions
Cutting without measuring is a waste of effort. Cut your #1 leak first.
30-Second Baseline (Do it once and you're done)
โ
Diet your CLAUDE.md to under 200 lines โ
Keep only 3 active MCPs โ
Extended thinking: Default OFF, use only when needed โ
.claudeignore: Exclude large generated files โ
Get into the habit of /clear when a task is finished
7 Advanced Tips with Huge Impact
โ Make Plan Mode the Default
Shift+Tab ร 2 before expensive tasks. Plan without touching code. Use this for broad requests like "Refactor this." It drastically reduces the ratio of tokens wasted on failed attempts.
โก Model Switching
80% Daily Coding โ Sonnet; Complex Reasoning โ Opus. Commands: /model sonnet, /model opus.
opusplan mode: Plan with Opus, implement with Sonnet. Can save 60% in costs.
โข Use Subagents Selectively
They run in a separate context and return only a summary to the main thread. Use them only for heavy explorationโfor small tasks, the overhead is actually larger. Rule: Only use when the main context savings > subagent start cost.
โฃ Active use of /compact
Waiting for the 80% context warning is too late. It will compress everything into noise.
Correct usage:
- At the end of every task phase
- Give a summary guide before calling
/compact: "Keep X, Y, Z and discard the rest."
โค Read with Precise File Ranges
โ "Look at the whole codebase" โ "Look at lines 50-120 of src/auth.js and improve error handling."
The difference is massive.
โฅ Session Handoff Notes
When a session gets long, before ending it:
"Summarize the work done so far, next steps, and important decisions in under 500 tokens."
Paste this into a new session = dozens of times fewer tokens than reconstructing the whole history.
โฆ Use Slash Commands for Repetitive Tasks
Don't explain frequent patterns (PR review formats, test rules) in natural language every time. Define them as Slash commands โ Deterministic and lightweight. Much more efficient than putting them in CLAUDE.md.
Common Pitfalls
โ "It's convenient to put everything in CLAUDE.md" โ You pay that cost every single turn. โ "Subagents are always cheaper" โ They are actually more expensive for small tasks. โ "The larger the context, the smarter it is" โ Opposite. Quality drops due to context rot. โ "Upgrading Pro to Max will solve it" โ The same inefficiencies will just cost 5x more. Fix the leaks first.
Token waste is a behavioral issue, not a limit issue.
Running /context once, dieting CLAUDE.md, cleaning up MCPs, and controlling Extended Thinking will solve most problems.
Remember that every message pays the cost of all previous messages, and you'll see exactly where to cut.
Prompts for Beginners
For Claude Code users (Self-diagnosis & Diet set)
Run the /context command and analyze the results.
Then, do the following in order:
- Tell me the top 3 items taking up the most tokens.
- For each, suggest a specific action I can take right now to reduce them (including estimated token savings).
- Read my CLAUDE.md and suggest a version dieted down to under 200 lines / 600 tokens. Recommend where to move the removed items (Skills? Slash commands? Or just delete?).
- Finally, check for other leaks like Extended thinking or MCP tool organization.
Since I'm a beginner, please organize the results by priority: "Do right now / Do when you have time."
For Claude.ai Chat users (Conversation Hygiene)
Copy and paste this when the conversation gets long, slow, or hits limits:
Summarize only the truly important information from this conversation in under 500 characters. Exclude all trial and error, tangents, and greetings. Focus only on core conclusions, decisions, and the next steps. I will copy this to start a new conversation, so organize it so I can pick up right where I left off.
Just using these two prompts will help you use AI in a much more pleasant environment without wasting tokens. If this helped, please give it a like!
If you have any other questions, leave them in the comments! :)





