MATINTELLECT

I thought tokens go to the model's answers. Turns out almost everything goes somewhere else

Yesterday I did a deep analysis and pinned down the five main factors, from strongest to weakest:

1. A long conversation in one window
The model doesn't remember the conversation between steps. On every message of yours and every action of its own it rereads the whole conversation from the very beginning. The longer the window, the more expensive each next step. By default, conversation compression only kicks in when the window is almost full. Until then, hundreds of steps run on a huge window

👉 Compress the conversation earlier (/compact) or start a new chat when you've finished a stage
💡 I cut the working window in settings from 1M to 500k tokens

2. Sub-agents
When Claude/Codex launches helpers for research, checks and running tasks, each one has its own window, and each one burns tokens too. Several helpers at once can eat more than the main conversation

👉 Launch sub-agents only for truly important tasks
💡 After Opus 5.5 came out this problem went away completely, and the model really only calls sub-agents where they're needed

3. Breaks longer than an hour
Anthropic/OpenAI keeps your conversation in fast memory for about an hour. After a long pause the whole window gets loaded again from scratch, and the bigger it is, the more it costs

👉 Before bed or a long break, compress the conversation with the command (/compact)
💡 For me this was a real eye-opener 😳 Now I never leave processes hanging without a reply for more than 1 hour

4. The conversation's starting weight
Before the first word, the conversation already weighs something: your rules, the skills list, connected add-ons, memory. All of it rides along with every step

👉 Turn off add-ons you don't use and keep your rules file short
💡 I audited everything that loads into context and cleaned out 25% of junk - Claude.md/Agents.md, Memory.md, MCP and so on

My example. I broke down my usage for the month:

• 97% of all tokens went to rereading old conversation, while the model's actual answers took less than 1% 🤯
• the window grew to almost a million tokens, and each step reread about 520 thousand on average
• helpers ate more than half of the usage in a week, 57%
• a new chat weighed about 80 thousand tokens before the first word

Main takeaway: you're not paying for the answers, you're paying for how many times the model rereads everything that came before them 😉

Share:

No comments yet

Leave a Comment

Fields marked with an asterisk (*) are required