YouMind
Войти

From Token Maxing to Token Management

@fukkyy
ЯПОНСКИЙ01 июн. 2026 г.
293K
606
108
0
726

Суть

As the era of unlimited AI usage ends, companies must shift from Token Maxing to Token Management to ensure ROI. This article explores how AI costs are becoming a critical management issue comparable to labor and marketing expenses.

Today, we discuss how the phase of "Token Maxing"—maximizing AI token usage—has come full circle, and the trend is shifting toward "Token Management," which questions the ROI of AI. Token management is now a management issue as critical as personnel planning or marketing investment decisions.

Overseas, this concept of "Token Management" is often discussed more technically as "Token Optimization." For the sake of clarity, I will use the term Token Management here, but please consider it synonymous with Token Optimization.

Token Maxing and Its Backlash

福島良典 | LayerX - inline image

Over the past year, the use of generative AI has been in a phase that could be called "Token Maxing." It was about showering tasks with tokens, hitting the strongest models repeatedly, and feeding in context to the absolute limit. AI vendors also ran with usage expansion as a success metric, effectively providing subsidies through flat-rate subscriptions. On the user side, using the strongest model without worrying about cost was talked about as a competitive advantage.

However, entering 2026, a clear backlash to this structure has begun to emerge.

Uber is a symbolic example. The company set up a Claude Code leaderboard internally to have engineers compete on usage, and as a result, they exhausted their 2026 AI coding budget in just four months. Uber COO Andrew Macdonald mentioned in a podcast:

"From now on, we must discuss token consumption and its associated costs alongside labor costs. If we can't draw a straight line showing that features reaching customers increase as much as we spend, this trade-off is difficult to justify."

https://www.youtube.com/watch?v=y_mQ6xLcKyc&t=1776s

Similar concerns are being voiced by those selling AI. Microsoft's Satya Nadella stated at the Morgan Stanley TMT Conference in March 2026:

"Why are we using so many tokens? In many cases, we are doing the worst things in terms of token efficiency. What everyone will work on in the next year is tool usage and software to make AI efficient."

https://www.investing.com/news/transcripts/microsoft-at-morgan-stanley-conference-ais-transformative-role-93CH-4542000

In fact, Microsoft has significantly scaled back the direct Claude Code licenses it had just introduced internally.

From here on, AI will not be a tool where you can just "use as much as you want," but a subject where ROI is questioned, just like labor or marketing costs. The point for management will shift from how many tokens were used to what those tokens produced and where to allocate them next.

Large-Scale Token Usage is a Prerequisite for Competition

While this sounds a bit negative, the premise is that the evolution of AI is extremely powerful. In the development field, there is no option not to use coding agents. Companies that cannot afford a certain level of token costs are essentially withdrawing from the competition. Moreover, the hurdle for that "certain level" of token cost is quite high. We must face the harsh reality that you cannot accelerate without investment capacity.

In fact, it is not uncommon for the token costs consumed by engineers to reach tens of percent or even half of their salary. It is becoming necessary to think of the "cost of hiring one engineer" not just as "salary," but as a set of "salary plus AI token costs." For management, this is the emergence of a completely new cost category.

Token Maxing was not a meaningless movement; by using AI at full throttle, I believe learning progressed rapidly regarding where to step on the gas and where to exercise more control. Moving forward, companies that can perform Token Management—stepping where they should and controlling where they should—will stand out.

Token Management has become a management issue itself. Token Maxing was an important first step for the structural transformation of companies.

At our company, the phase of just using tokens for the sake of it has ended, and we are in a phase of how to produce outcomes. Within LayerX's "Compass" (guiding principles), there are directives such as "Make AI a growth driver" and "Don't build things that won't be used. Don't fall into the AI build trap."

福島良典 | LayerX - inline image

From https://speakerdeck.com/layerx/compass_202209?slide=34

福島良典 | LayerX - inline image

From https://speakerdeck.com/layerx/compass_202209?slide=36

As a result of these guidelines, we optimized tokens internally. We used them where they should be used and optimized where they should be optimized, yet the number of tokens used has not changed significantly from when we were doing Token Maxing at LayerX. I feel that large-scale token usage above a certain level has become a prerequisite for competition.

Coding Agents are Merely a "Precedent"

What's important here is that this explosive token consumption is not limited to the specialized field of coding.

Currently, all kinds of business operations besides coding are being solved in the same way as coding agents. Signal collection and proposal creation for sales, customer support inquiries, automatic expense reporting and approval, legal contract reviews, marketing ad operations, HR recruitment screening—these are being redesigned one after another as autonomous execution agents.

Autonomous agents consume tokens in quantities incomparable to the previous "one-question, one-answer" way humans used them. As agents do more work, all AI usage will converge into the same pattern we see now with coding agents. Coding agents just happened to be the first precedent to reach that pattern.

福島良典 | LayerX - inline image

The AI token cost management challenge, which currently looks like an "engineer-only problem," will soon become a company-wide management issue. This is because every business area within a company will have the same token consumption profile within the next few years.

Token Costs are a Time Bomb

福島良典 | LayerX - inline image

There is another fact we must recognize here. The AI subscription fees we are currently paying are significantly lower than the actual costs.

Currently, it is common for foundation model providers to offer AI subscriptions at a loss, and there is a huge gap between the flat fees we pay and the actual costs incurred. If you recalculate the same usage with API pay-as-you-go pricing, the cost jumps by orders of magnitude. The subscription format makes the actual cost invisible to the user.

The problem is that foundation model companies could make this difference in actual cost apparent at any time. The moment the subsidized portion starts being passed on to customers in the form of a shift from flat-rate to token-based billing, guidance to higher plans, or price hikes, the sense of "tens of thousands of yen per ID per month, costing a company of 100 employees tens of millions of yen annually" could jump to the order of "hundreds of millions or billions of yen annually for the whole company" in a short period. The difference in actual cost hidden by subscriptions becomes the power of the time bomb.

The first thing needed is to make what is happening inside the company visible before the explosion surfaces.

(This article is helpful for details: https://www.thestateofbrand.com/news/ai-subscription-time-bomb

First, Visualization

The first thing needed for this problem is something extremely basic: making it visible "who is using which model and how many tokens." That's it.

For people, we already have vast management foundations like personnel evaluations, goal management, salary systems, and organizational design. For marketing expenses, it's natural to optimize while measuring CAC and payback. Procurement has specialized teams and workflows.

However, for AI, even this most basic visualization is not yet in place. For CEOs and IT departments, the fact that they cannot see "who is spending how much on AI right now" is already a major stressor. Value is created just by putting numbers on a dashboard.

The "AI Token Advisor" we are currently providing is a product exactly designed to fill this gap. It is a simple tool to visualize who used which model for what task and how many tokens. (Bakuraku users can use one ID for free!)

福島良典 | LayerX - inline image

From https://bakuraku.jp/resources/product/pre-registration-ai-token-advisor/

However, visualization is only the entrance. What's truly important is the question that inevitably arises beyond that.

"Should that task really have used that model?"

As visualization progresses, the next question that always arises is: "Should that task really have used that model?"

For example, using a top-tier reasoning model (like Claude Opus 4.8 or GPT-5.5; the latest models as of May 2026) for a question like "What's the weather today?" is clearly excessive. The lightest model is sufficient, and it might not even need to be a reasoning model. Similarly, there are massive cases within companies where expensive models continue to be used for mature tasks like routine email creation.

Trends also vary by person. People who are not used to using AI tend to just pick the strongest model for now. Conversely, those who use AI deeply switch models depending on the nature of the task. This difference, in terms of both cost and quality, becomes too large for an organization to ignore.

What's important here is the fact that all input to AI is ultimately converted into the form of a "prompt." Chat questions and context passed to agents are all internally input to the model as prompts. In other words, if you look at the prompt, you can pretty much tell what that person is making the AI do.

Starting from this prompt, we can detect mismatches between tasks and models and provide guidance like, "For this task, a lighter model is sufficient." Going further, instead of people choosing models every time, the system will automatically allocate the optimal model. We are moving toward such a world.

福島良典 | LayerX - inline image

In the age of agents, this becomes even more important. Since agents execute autonomously, there is no room for a human to think every time, "Is it a luxury to use Opus now?" From the start, there will be no choice but to leave the optimal model selection based on the nature of the task to the system.

The services we provide at Bakuraku will be offered as managed services, including the backend implementation that selects the optimal model without the user having to be conscious of it. AI Token Advisor is merely the entrance.

Summary

AI token costs are already exceeding the scale where they can be handled with a "just throw money at it" approach. For some companies, they are becoming expenses on par with labor and marketing costs.

Marketing costs are strictly managed by CAC, and no one operates by saying, "Just run ads regardless of CAC." The same goes for labor costs; salary, hiring, and placement are all managed. You don't just hire people at any salary.

On this same level, AI costs will enter the scope of management. The first thing needed then is to visualize "how much is being spent." And next is to check "whether the appropriate model is being chosen for the task and the person."

It may look unglamorous, but mastering these two points will become a critical management issue in the coming years.

For those who want to implement such a future together, LayerX is actively recruiting AI Builders!

福島良典 | LayerX - inline image

We are actively hiring! https://jobs.layerx.co.jp/

Сохранение в один клик

Используйте YouMind для глубокого чтения вирусных статей с помощью ИИ

Сохраняйте источники, задавайте точные вопросы, обобщайте аргументы и превращайте вирусные статьи в полезные заметки в одном рабочем пространстве ИИ.

Исследовать YouMind
Для авторов

Превратите ваш Markdown в аккуратную статью для 𝕏

Когда вы публикуете длинные тексты, изображения, таблицы и блоки кода, форматирование в 𝕏 становится мучением. YouMind превращает полный черновик в Markdown в чистую статью, готовую к публикации в 𝕏.

Попробовать Markdown для 𝕏

Другие паттерны для анализа

Недавние виральные статьи

Смотреть другие виральные статьи