If you use Claude Code on a subscription, you live inside two windows: a five-hour one and a seven-day one. You usually find out where they stand when you hit the wall.
I wanted that answer in front of me the whole time, so I built burnrate: a small mod that puts one row above the Claude Code prompt and tells you how fast your plan windows are filling, when they will run out, and when your prompt cache was thrown away.

Live page: nemke82.github.io/claude-code-burnrate · Code: github.com/nemke82/claude-code-burnrate
What the row tells you
● burnrate 5h 23% +4%/h 7d 41% 98% cached
● burnrate 5h 62% +30%/h full in 1h 16m 7d 41% 98% cached
✖ burnrate 5h 24% +4%/h 7d 41% 0% cached miss: model changed, ~81k rewritten
Reading left to right:
5h 62%is the plan window that matters most right now. It turns yellow at 70% and red at 90%.+30%/his the pace: how fast that window has been filling over the last half hour.full in 1h 16mappears only when, at that pace, the window fills before it resets. That is the one warning worth interrupting you for.98% cachedis how much of the last request was served from the prompt cache.miss: model changed, ~81k rewrittenshows up when a request wrote again what the previous one had cached.
On a narrow terminal the row drops detail in order of importance. The window and its warning go last.
/burn for the detail
Type /burn and a pane opens with the session's totals, every cache miss, both plan windows with their reset times and pace, and the last requests one by one.

Both animations are generated from the mod's own rendering code with a scripted session, so every line is what burnrate draws for those numbers. Here is what it showed in the real session I built it in:
│ Session 8 requests · 99% cached · 8.5k output │
│ 5h ███░░░░░░░░░░░░░░░░░ 15% resets in 4h 6m │
│ 7d ██░░░░░░░░░░░░░░░░░░ 10% resets in 3d 20h │
● burnrate 5h 15% 7d 10% 100% cached
Why cache misses belong on a quota meter
Claude Code caches the front of your conversation, so each new request only pays full price for what is new. When that cache is lost, the whole context is written again, and that counts against your window.
A miss is easy to cause without noticing. burnrate reports what it can tell about each one:
| Shown as | What it means |
| --- | --- |
| model changed | the request went to a different model than the one before |
| after 12m idle | more than five minutes passed since the previous response |
| likely prefix change | neither of the above; usually effort, tools, the system prompt or CLAUDE.md changed |
What it does not claim
I would rather the meter say less than say something wrong, so a few things are deliberate:
- No dollars. The only cost figure available is list price. That is not what a subscription pays, nor a custom contract, so burnrate shows quota and tokens, which are true for everyone.
- Token counts for a miss are estimates, marked with
~. The API reports how many tokens were read and written, not which ones. - A pause is not called an expiry. The mod cannot see whether a request asked for the 5-minute or the 1-hour cache, so it reports the pause and leaves the conclusion to you.
- The pace is a projection, not a forecast. One heavy turn moves it, and the windows belong to your account, so other sessions move them too.
- One point is not a trend. A window that moved a single point does not raise a "full in" warning.
How it works
A Claude Code mod is a small TypeScript module of hooks. burnrate uses four:
turn.stepsees every model request and its usage: tokens read from cache, written to cache, sent uncached.session.measuredelivers the plan windows whenever they move.ui.renderdraws the row above the prompt and the/burnpane.command.runanswers/burn.
Everything that decides what you see is a pure function with no access to the engine: meter.ts detects misses, pace.ts works out the pace, view.ts lays out the row and the pane as plain data. The hooks only feed them and draw the result.
Tested, including the awkward cases
That split makes the logic easy to test. There are 44 tests; these are the ones for the pace alone:
tests/pace.test.ts:
(pass) a reading is added only when the window moved
(pass) a share that fell is a reset: the window starts its history over, the others keep theirs
(pass) readings past the kept history are dropped
(pass) no pace without history, nor from too little of it: five minutes, an hour for the slower windows
(pass) a rise of one point is not enough for a warning
(pass) a pace that fills the window before it resets
(pass) a pace the reset beats
(pass) the pace falls off while nothing is sent
(pass) without a reset time there is a rate and a time to full, but no warning
(pass) a full window has no time to full
(pass) fmtPace: per hour, per day for the seven-day window, one decimal under ten
44 pass
0 fail
Ran 44 tests across 4 files.
A test reads like the sentence it proves:
test('a pace that fills the window before it resets', () => {
// 47% to 62% in half an hour: 30 points an hour, 38 to go, three hours to the reset
const pace = paceOf([at(-30, 47)], five(62), T0)
expect(pace?.perHour).toBe(30)
expect(Math.round((pace?.fullInMs ?? 0) / MIN)).toBe(76)
expect(pace?.hitsLimit).toBe(true)
})
The row and the pane are tested too, by mounting them on every surface Claude Code draws to (terminal, desktop, VS Code, mobile) and checking what comes out.
Built in a morning, with a second opinion
I built burnrate in one session with Claude Code, and had a Codex agent in the pane next to it review the work. That review changed the product more than once:
- It blocked the commit because my "tokens rewritten" figure could exceed what was actually lost. The estimate is now capped and labelled as one.
- It caught that I measured idle time from the start of the previous request, so a slow response looked like a pause.
- It first told me to lead with dollars, then reversed itself when I pointed out most users are on quotas. Dollars came out entirely.
The demo animation caught one more: a one-point move on the seven-day window produced a confident "full in 3 days" warning. That is why a single point no longer counts.
Try it
/plugin marketplace add nemke82/claude-code-burnrate
/plugin install burnrate@burnrate
Mods are early access, so start Claude Code with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (version 2.1.259 or later). The mod API may change between releases, and burnrate is young: the thresholds have not been tuned against many real sessions yet.
What's next
- Tuning against real sessions. The thresholds (what counts as a miss, when a warning fires) are reasoned and tested, but not yet tuned on many real sessions.
- A few options. The warning thresholds and whether the row shows at all should be yours to set.
- Subagents. Only the main conversation is counted today; subagents have caches of their own and are left out.
Tell me what your row shows
The most useful thing you can send me is a case where burnrate got it wrong: a "miss" that was not one, a warning that fired too early, a pace that made no sense. Open an issue with what the row or /burn showed and what you were doing, and I will use it to tune the thresholds.
Page: nemke82.github.io/claude-code-burnrate · Source, tests and the demo scripts: github.com/nemke82/claude-code-burnrate. Issues and pull requests are welcome.