YouMind
Sign in

Achieve a 67% Token Reduction: The "Escalation" Strategy for Claude Code

@ClaudeCode_love
JAPANESEMay 04, 2026
366K
551
47
4
1.3K

TL;DR

This guide outlines a viral strategy for optimizing Claude Code usage, focusing on separating planning from execution, using external memory files, and strategically escalating between Haiku, Sonnet, and Opus models to avoid rate limits.

"Ugh, I hit the Claude Code usage limit again! 😭 So stingy! 💢" I get the feeling. But maybe it's your operation method that's the problem? → So what should you do? → Read this article → Understand token-saving methods → Problem solved for everyone!!!!

Let's get into it!!!

Have you ever experienced this while using Claude Code?

Claude Code Studio - inline image

・Suddenly seeing "Usage limit reached" in the middle of a prompt

・Hitting rate limits every few hours despite being on a $200/month plan

・Losing focus and productivity because you're worried about limits

・Worrying every month about whether you should upgrade your plan to avoid limits

・Stopping in the middle of important work and ending up running to another AI

An article by Miles Deutscher (@milesdeutscher), a top AI influencer with 670,000 followers overseas, is currently going viral with 3.35 million likes 😳

Claude Code Studio - inline image

He himself was hitting rate limits daily while using the $200/month Anthropic plan. However, by "re-understanding the fundamental mechanism of Claude," he hasn't hit a token limit once in the past three weeks.

Today, I'll break down those contents in an easy-to-understand way 👇

Original post here: https://x.com/milesdeutscher/status/2049618781841031551

■ 𝗦𝘁𝗲𝗽 𝟭: 𝗣𝗹𝗮𝗻𝗻𝗶𝗻𝗴 (Completely separate planning and execution)

Claude Code Studio - inline image

Miles first points out: "Don't brainstorm with Claude Opus."

Many people probably do this. You get an idea, throw it at Opus to bounce ideas off it. Before you know it, 30 minutes have passed, and you've reached the limit. Sound familiar?

The fact Miles discovered through deep diving is this:

"Text chat itself doesn't consume that many tokens. What really consumes them are execution-type tasks like coding, building, and designing."

In other words, just by clearly separating the phase of thinking about what to make (Planning) from the phase of actually making it (Execution), you can drastically reduce the consumption of high-cost models.

Miles provides a specific comparison. In the case of two people making the same finance tracking app:

Claude Code Studio - inline image

Person A: Spends only 2 minutes planning and starts building with a weak design. Result: 3 re-dos.

Person B: Spends 20 minutes planning to solidify the design and completes the build in 1 go.

Person B saved about 67% of tokens on this task alone. That's a $1.50 difference in cost. Considering there are many tasks in a day, it becomes a difference of dozens of dollars per month.

For those using Claude Code, the "Plan Mode" entered by pressing Shift+Tab×2 is exactly the feature that embodies this philosophy.

Claude Code Studio - inline image

In Plan Mode, Claude focuses on design and planning without writing code. This means you can solidify the architecture and policy without consuming execution tokens.

Furthermore, Miles' style is to leave the planning phase itself to cheaper models. Instead of bouncing ideas off Opus, Haiku is sufficient. Haiku is smart enough for brainstorming, and the cost is orders of magnitude cheaper.

Practice points:

・Do ideation, brainstorming, and design with Haiku

・Switch to Opus only after the design is solid and you're "ready to build"

・Get into the habit of using Plan Mode (Shift+Tab×2) every time in Claude Code

・The more you skimp on "thinking time," the more "re-dos" increase, leading to a total loss

■ 𝗦𝘁𝗲𝗽 𝟮: 𝗖𝗵𝗮𝘁 𝗟𝗲𝗻𝗴𝘁𝗵 (Chat length rules everything)

Claude Code Studio - inline image

Miles says long chats are silent killers. This is the biggest pitfall many people overlook.

The mechanism is this: Every time you send a message, Claude re-reads the entire context within that chat. That means:

Claude Code Studio - inline image

・When the chat is 10 messages: It reads 10 messages' worth of tokens

・When the chat is 100 messages: It reads 100 messages' worth of tokens

As the chat gets longer, the cost per message increases exponentially. And cost isn't the only problem. As old information gets mixed in, the quality of Claude's output itself degrades. It gets pulled by irrelevant past context, and off-target answers increase.

Miles has two solutions.

𝟭. Utilize 𝗣𝗿𝗼𝗷𝗲𝗰𝘁𝘀

Claude Code Studio - inline image

If you do the same type of task repeatedly, create multiple sub-chats within a Project instead of one long chat.

Miles himself has a Project for writing on X and opens a new chat every time he writes a new article. Since Project settings (Instructions) are shared across all chats, there's no need to re-explain "I am this kind of person, write in this style" every time.

Even smarter is to include this sentence in the Project Instructions:

"Be cognisant of the fact I'm trying to save account usage. Be concise in your answers, and when appropriate, advise me on when I should start a new chat or any other tips that may help me reduce token usage."

With just this, Claude itself becomes a token-saving advisor. It will start telling you, "It's probably time to move to a new chat."

𝟮. Compressed context transfer with Mega Prompts

Claude Code Studio - inline image

If you absolutely want to carry over the context of the current chat to the next one, say this at the end of the chat:

"I'm moving to a new chat; give me a prompt I can use to restart this session without losing any of our context from this conversation."

Claude will generate a single prompt that compresses the entire context. Just paste this at the beginning of a new chat to restart with a lightweight chat without context loss.

The golden rule to remember:

Claude Code Studio - inline image

"Three short chats" are overwhelmingly more token-efficient than "one ultra-long chat." If in doubt, open a new chat. This alone will drastically reduce the frequency of hitting limits.

■ 𝗦𝘁𝗲𝗽 𝟯: 𝗣𝗿𝗼𝗽𝗲𝗿 𝗠𝗲𝗺𝗼𝗿𝘆 (Persist Claude's memory in external files)

Claude Code Studio - inline image

One of Claude's biggest weaknesses is that it forgets context.

By default, Claude remembers almost none of your preferences or past instructions. As a result, what happens is:

・Explaining the same prerequisites every time → Consuming tokens for that

・Repeating mistakes that were corrected in the past → Consuming tokens in the interaction to correct them again

・Forgetting preferences and giving unnecessary output → Consuming tokens for retakes

Miles introduces a way to fundamentally break this vicious cycle.

The method is simple. Create a folder on your desktop and place two Markdown files inside.

Claude Code Studio - inline image

𝗜𝗻𝘀𝘁𝗿𝘂𝗰𝘁𝗶𝗼𝗻𝘀.𝗠𝗗 (Instruction Sheet)

A file to write permanent rules and instructions for Claude.

Example structure:

・## Who you are → Your role/expertise

・## What you do → Behavior expected of Claude

・## Rules → Rules you want it to strictly follow

And put the most important line here:

"Update Memory.MD with my preferences over time."

With this instruction, Claude will automatically write your preferences and corrections learned during the conversation into the second file.

𝗠𝗲𝗺𝗼𝗿𝘆.𝗠𝗗 (Memory File)

A file that functions as Claude's "second brain." It gets smarter the more you use it.

Example structure:

・## Preferences → Preferred styles, formats

・## Corrections → Matters corrected in the past

・## Patterns → Patterns used repeatedly

Specific example: If you say "don't use em dashes" once, Claude records it in this file. From the next time on, em dashes won't appear even if you say nothing. If you say "use ■ instead of # for headings," that will also be recorded.

Claude Code Studio - inline image

Just attach this folder to Claude Code/Cowork to complete the setup. Since Claude reads the contents of the folder every time, context is maintained across chats.

Miles says once you start using it, you can't go back. The fact that tokens spent on re-explanation become zero is quite significant in terms of experience.

■ 𝗦𝘁𝗲𝗽 𝟰: 𝗠𝗼𝗱𝗲𝗹 𝗦𝘁𝗮𝗰𝗸𝗶𝗻𝗴 & 𝗦𝗲𝗹𝗲𝗰𝘁𝗶𝗼𝗻 (Save 90% by using models appropriately)

"Using Opus 4.7 for everything is a complete waste," Miles asserts.

Claude Code Studio - inline image

A common mistake people make is thinking, "I'll be fine if I always use the smartest model." But this is like "taking a Ferrari to the local convenience store."

Miles practices the "Escalation Method."

Claude Code Studio - inline image

Haiku (light tasks) → Sonnet (medium tasks) → Opus (heavy tasks/final finishing)

Start in this order and switch to a higher model only when the capability is truly insufficient. In his experience, 90% of tasks can be handled sufficiently by models other than Opus, and Opus is only truly needed for the remaining 10%.

Further fine-tuning:

Claude Code Studio - inline image

・𝗘𝘅𝘁𝗲𝗻𝗱𝗲𝗱 𝗧𝗵𝗶𝗻𝗸𝗶𝗻𝗴: Keep it off normally. Turn it on only for complex reasoning or mathematical tasks. When on, token consumption jumps, so use it only when truly necessary.

・𝗦𝘁𝘆𝗹𝗲𝘀 (Style Settings): You can switch to the "Concise" style from Claude's home screen. This alone makes answers short and simple, significantly reducing output tokens. Many people don't even know this feature exists.

・𝗟𝗼𝘄 𝗘𝗳𝗳𝗼𝗿𝘁: In Claude Code, you can select "Low" effort mode. This is sufficient for simple tasks and increases processing speed.

And don't forget options other than Claude. For simple tasks like news search, research, and summarization, free or cheap open-source models like Kimi or DeepSeek are sufficient. Save Claude's quota for "things only Claude can do."

■ 𝗦𝘁𝗲𝗽 𝟱: 𝗧𝗼𝗼𝗹 𝗦𝗽𝗹𝗶𝘁𝘁𝗶𝗻𝗴 (Strategically use quotas for each tool)

Claude Code Studio - inline image

A fact most people haven't noticed: each Claude tool has its own independent usage parameters.

Specifically:

Claude Code Studio - inline image

・Claude Code / Claude Chat → Share the same plan's usage quota

・Claude Design → Completely separate quota

If you don't know this mechanism, what happens? For example, you have Claude Code create a UI design mockup. This consumes the Code/Chat quota. But the separate tool, Claude Design, has its unused quota completely remaining. If you do the same design task in Claude Design, you can avoid consuming the Code/Chat quota at all.

It's most cost-effective to use each tool for its originally designed purpose.

Miles' rules:

・Coding → Claude Code

・Design → Claude Design

・Dialogue/Analysis → Claude Chat

・Use each tool for what it's good at, and don't force it to do what it's not.

■ 𝗕𝗼𝗻𝘂𝘀 𝗧𝗶𝗽𝘀 (Collection of additional techniques you can use immediately)

Claude Code Studio - inline image

・Purchase additional credits: Before considering a plan upgrade like $20→$100, there's an option to buy just a few dollars' worth of additional credits. This is enough when you're a little short at the end of the month.

・Claude Skills: Build skills to automate repetitive tasks. Instead of explaining the same procedure every time, save it as a skill to execute with one command.

・Usage Tracking: Get into the habit of checking usage status regularly. In Claude Code, you can check immediately with the /Usage command. If you know "what % is left," you can adjust how you use it.

・Overview Section: A newly added feature where you can see a dashboard with an overview of usage status at a glance.

・Change behavior when approaching limits: When less than 20% remains, consciously switch modes by switching to Haiku, turning off Extended Thinking, keeping chats short, etc.

■ Summary: Zero limits achieved for 3 weeks with this method

Claude Code Studio - inline image

Miles says he hasn't hit a token limit once in the three weeks since practicing these 5 steps. Without changing his $200/month plan.

To organize the points:

Claude Code Studio - inline image

・Step 1: Planning with Haiku, execution with Opus. 67% reduction just by separating phases.

・Step 2: Keep chats short and manage with Projects. 3 short chats > 1 long chat.

・Step 3: Externalize memory with Memory.MD to zero out re-explanation costs.

・Step 4: Use the escalation method to send 90% to models other than Opus. Also utilize Styles and Effort settings.

・Step 5: Understand the difference in usage quotas for each tool and use the right tool for the right job.

Honestly, the prospect of AI usage costs getting cheaper in the future is slim. Rather, as models become more high-performance, token unit prices tend to rise. That's why learning the "correct way to use" now directly leads to long-term savings.

As Miles says, the problem isn't that the "plan is cheap," but that the "usage is wrong." If used correctly, a life without hitting limits on your current plan is entirely achievable.

For those who found this article even slightly helpful.

Claude Code Studio - inline image

𝗖𝗹𝗮𝘂𝗱𝗲 𝗖𝗼𝗱𝗲 𝗦𝘁𝘂𝗱𝗶𝗼 @ 𝗝𝗮𝗽𝗮𝗻 (@ClaudeCode_love) is an account run by three Claude Code enthusiasts.

We post daily about practical CLI utilization and automation.

Currently co-developing an AI agent with a listed company.

Our usual posts 👇

・Real product development examples using Claude Code and Claude

・Organization of Claude Code utilization / Vibe Coding / development trends

・Latest information on Claude Code from overseas

From development philosophy to design, implementation, and improvement,

we summarize overseas and primary information to get working products out into the world, not just "finish making them."

If you're interested, please follow and check it out 👀 I think it'll be beneficial!

One-click save

Use YouMind for AI deep reading of viral articles

Save the source, ask focused questions, summarize the argument, and turn a viral article into reusable notes in one AI workspace.

Explore YouMind
For creators

Turn your Markdown into a clean 𝕏 article

When you publish your own long-form writing, images, tables, and code blocks make 𝕏 formatting painful. YouMind turns a full Markdown draft into a clean, ready-to-post 𝕏 article.

Try Markdown to 𝕏

More patterns to decode

Recent viral articles

Explore more viral articles