Where do your AI tokens go?
We typed 11 words and the agent used 150,194 tokens. About 99% went to re-reading, and the answer itself was the small part. Here is why, and the habits that keep it lean.
We asked an AI agent (Codex) for a quote in 11 words:
Quote Cafe Luma for a Website Refresh. Summer discount. Send it.
It was the same job as our last note: the client and the data are fictional, the run and its log are real. The agent read 150,194 tokens and wrote 1,930. When we answered "yes" so it could send the email, that one word re-read 199,357 tokens.
What a token is
Models read and write in pieces of words called tokens. OpenAI's rule of thumb for English is that 100 tokens are about 75 words. So 150,194 tokens are around 113,000 words, about the length of a whole novel.
Where they went
A model has no memory between steps. An agent works in steps: read a file, run a command, check the result. At every step it sends everything again: the instructions file (CLAUDE.md or AGENTS.md), the whole conversation, every file it has opened and every result so far.
In our three logged runs, about 99% of the tokens went to that re-reading (98.7%, 98.9% and 99.0%). Writing the answer was about 1%.
More context gives better answers, so the goal isn't an empty context. The goal is a context with nothing it doesn't need. Part of the re-reading is cached, which makes it cheaper: 80% of the "yes" was served from cache. Cheaper, but still counted.
What we do now
- Save it, start fresh. When a task is done, we save the decisions, where we stopped and the next steps in a one-page file, and open a new chat that reads only that page.
- Right model, right effort. A big model with high effort to plan and decide, a small one with low effort for renaming, sorting or summarising. In Claude Code,
/modeland/effortchange this. - Divide and conquer, only for big jobs. Helpers (subagents) can run on smaller models and return a short summary, so the main conversation stays light. Each helper starts from zero, though, so a small job costs more when you split it.
- Keep CLAUDE.md or AGENTS.md short. It is loaded into every session. Claude Code's docs suggest under 200 lines, and moving rarely used instructions into skills, which load only when needed.
- Name the exact file. "Look around" makes the agent search and read.
- Give it a ready path. The same 11-word request took 10 steps and 150,194 tokens in one run, and 4 steps and 71,258 tokens in another where the agent had short instructions, the data and a script ready. The setups were different, so this isn't a lab test, but fewer steps meant fewer re-reads.
- Mind long breaks. The cache doesn't last forever. In Claude Code on a subscription it lasts an hour; after that, the whole conversation is processed again. A new task is a good moment for
/clear, which costs nothing.
Next up on Instagram: eight of these habits in a carousel, ready to save.
Numbers from our own Codex logs, with a fictional client and data. Different model and CLI versions between runs, so not a controlled experiment. Commands and limits change by tool and plan. Sources: Claude Code docs, "Manage costs"; OpenAI Help Center, "What are tokens and how to count them?".