Developer Tracks 430 Hours of Claude Code Use, Finds 73% Waste
Mnimiy reports logging 430 hours and $1,340 in API spend across 90 days of Claude Code sessions, concluding that 73% of tokens went to nine overhead patterns. The article follows Anthropic's acknowledged usage-limit problems and offers a fix for each pattern.
Original post · 12 min read
430 hours of work. 6 million input tokens. $1,340 in API spend.
Then I sat down with the data and asked the only question that mattered: how much of this was actually doing my work, and how much was overhead?
The answer was uncomfortable: 73% of my tokens went to nine invisible patterns that I'd been doing on autopilot.
Not bad prompts, I write decent prompts.
Not big models when I needed small ones, I knew that one already.
Patterns deeper than that. Patterns nobody talks about because they're invisible until you instrument them.
If you're hitting Claude Max usage limits more than once a week, you have at least 4 of these. Probably 7.
Below: each pattern, exactly how much it costs, and the 30-second fix.
Why this matters now
Anthropic admitted the problem in late March 2026: "people are hitting usage limits in Claude Code way faster than expected."
Max 5 subscribers reported quota exhaustion in 19 minutes instead of the expected 5 hours.
The Pro $200/year users said the limit maxed out every Monday and didn't reset until Saturday - "out of 30 days I get to use Claude 12."
Anthropic acknowledged the issue. The peak-hours quota change in late March explained part of it. The rest was a prompt-caching bug — two independent bugs in the cache layer that silently inflated costs by 10-20x for some sessions.
A user reverse-engineered the Claude Code binary to find it (GitHub issue #40524). Downgrading to v2.1.34 helped some people.
Some of it is real. Most of it is you.
I'm not going to tell you to "use Haiku for simple tasks" or "start a new chat every 15 messages."
Below: the patterns that the obvious advice misses.
The methodology
I ran an HTTP proxy between Claude Code and the Anthropic API. Logged every request: full payload, response, token counts (input/output/cache), latency, model. 90 days. 430 hours of active work.
For each request, I categorized the tokens into:
Productive — content that directly informed my actual question
Cache hit (free) — system prompt + CLAUDE.md cached, no marginal cost
Cache miss (paid) — same content, recomputed because cache expired
Conversation history re-read — re-tokenizing previous messages on every turn
Hook injection — pre-pended context from PreToolUse / UserPromptSubmit hooks
Skill loading — skill SKILL.md content loaded into context for invocations
Tool use overhead — JSON schemas, tool definitions, tool result blocks
Extended thinking — <thinking> blocks
CLAUDE.md — project rules loaded every turn
Productive tokens: 27%. The other 73% is the 9 patterns below, ranked by how much they cost me.
Pattern 1. CLAUDE.md bloat (~14% of total tokens)
The pattern: my CLAUDE.md grew to 4,800 tokens over 6 months. Every turn loaded all 4,800 tokens. Every session loaded them again on each new request. Most of the rules were never relevant to the task at hand.
A 5,000-token CLAUDE.md costs you 5,000 tokens before you've typed a word. Every turn. Every session. A constant baseline tax. Multiply by 200 turns per week and it's 1 million tokens a week of your CLAUDE.md alone.
The 30-second fix:
If you're over 1,500 tokens combined, refactor:
Move framework-specific rules to project-level CLAUDE.md (only loads in that project)
Extract repeated patterns into skills (loaded only when invoked)
Delete anything you can't remember writing
Convert "explain why" verbose rules into 3-word imperatives
I cut mine from 4,800 to 900 tokens. Same behavior. 31% reduction in baseline cost, instantly.
Pattern 2. Conversation history re-reads (~13% of total tokens)
The pattern: every follow-up message re-tokenizes the entire conversation history. By message 30 in a chat, each turn is paying for messages 1-29 to be read again. The math: at ~500 tokens per exchange, message 30 costs 30× message 1.
I had sessions with 60+ messages. The last message was costing 60× the first. Tokens spent re-reading old context: catastrophic.
The 30-second fix:
Edit the prior message instead of follow-up. Up-arrow -> edit -> re-send. The bad exchange gets replaced, not stacked.
Hard cap conversations at 20 messages. When you cross 20, ask Claude to summarize what's been done and start a fresh chat with that summary as the first message.
Use /compact instead of /clear when you need continuity. /compact summarizes and restarts. /clear nukes everything.
I went from 60-message sessions to 15-message average. 40% drop in conversation re-read cost.
Pattern 3. Hook injection waste (~11% of total tokens)
The pattern: I had 4 plugins installed. Three of them registered UserPromptSubmit hooks that injected context. Combined: 6,200 tokens of hook injection on every prompt I submitted, before Claude even read what I asked.
These hooks are designed to be helpful. They inject branch names, recent file changes, instinct summaries, memory snippets. Each one is small. Together they're a wall.
The 30-second fix:
Aud… continue on X ↗