Guides
Where the tokens of an AI coding plan go, and how to cut the waste. The numbers come from real Claude Code use.
How Overplus measures a saving, and what it touchesWhere the numbers come from, what is real data and what is an estimate, what Overplus shows on Codex, and what stays on your Mac.Why you hit the Claude Code usage limit, and what uses it upA real week of Claude Code usage, broken down by token type. 92% of it was context re-reads and re-saves, not answers. Here is what that means for your limit.How to reduce Claude Code token usage: 9 settings that workThe Claude Code settings and environment variables that cut token use: compaction window, MCP and Bash output caps, subagent model, effort, deferred MCP tools and more.Opus or Sonnet in Claude Code? Model routing, effort and the prompt cacheWhen a cheaper model saves tokens in Claude Code and when it costs more. How effort levels work, and why switching models mid-chat rewrites your cache.What cuts LLM and coding-agent token useContext, thinking, model switches and tools: where the tokens of an LLM app or coding agent go, and which kinds of change cut them.