One setting cut my Claude Code cost per request by 39%. My first three mods: under 1%.

When mods landed in Claude Code, I wrote three of them. One of them, ci-watch, waits for CI in the background instead of letting the agent poll gh run list every few seconds. I was sure that one would show up on the bill.

Then I actually counted. Claude Code keeps every session as a .jsonl file under ~/.claude/projects, and every model request in it carries its token usage. So I took all of September and replayed it: where does the cost actually go, and what would the mods have saved?

The three mods together: under 1%. One line in settings.json cut the cost of a request by 39%. The cost is mostly about one thing, how much context every turn carries, and my mods don't touch that.

Before the numbers, the fine print: this is one person's month, and the 39% compares 15 days before the change with 7 days after, with different work in each. Treat it as a direction, not a benchmark.

How I counted

Raw token counts lie a bit: a token read from the prompt cache costs a tenth of a fresh one. So I weighted every request by API prices, relative to plain input:

  • input: 1x
  • cache read: 0.1x
  • cache write: 2x for the 1-hour cache, 1.25x for the 5-minute one
  • output: 5x

I'm on a Claude subscription, so this isn't what I paid. It's what the same work would cost at API prices, which is a fair way to see where the tokens go.

Requests show up more than once in a session file (one line per content block), so dedupe by requestId. Subagents live in their own files under subagents/. The core of it:

1import glob, json, os
2
3def cost(u):
4    cc = u.get("cache_creation") or {}
5    return (u.get("input_tokens", 0) + 0.1 * u.get("cache_read_input_tokens", 0)
6            + 2 * cc.get("ephemeral_1h_input_tokens", 0) + 1.25 * cc.get("ephemeral_5m_input_tokens", 0)
7            + 5 * u.get("output_tokens", 0))
8
9seen, total = set(), 0.0
10for path in glob.glob(os.path.expanduser("~/.claude/projects/**/*.jsonl"), recursive=True):
11    for line in open(path):
12        d = json.loads(line)
13        usage = (d.get("message") or {}).get("usage")
14        if usage and d.get("requestId") not in seen:
15            seen.add(d.get("requestId"))
16            total += cost(usage)
17

Everything below is a share of that total. I'm keeping absolute numbers out on purpose, the shape is what matters.

Where the tokens actually go

About 73% of the cost was cache reads. Cache writes were 20%, output 7%.

That's the important one. Every turn sends the whole conversation again. Most of it comes from the cache and is cheap per token, but you pay for it on every turn. So the cost is roughly context size times the number of turns, and what the agent writes barely matters.

And it's very uneven. Requests with more than 200k tokens of context were 31% of my requests and 54% of the cost.

Share of requests and share of cost by context size: under 50k 8% and 4%, 50k to 100k 27% and 14%, 100k to 200k 33% and 27%, over 200k 31% and 54%.

One setting: compact earlier

Claude Code compacts a session when the context gets close to the window. On a model with a 1M window that's at about 967k tokens by default, and every turn on the way there re-reads all of it.

There's a setting for that:

1{
2  "autoCompactWindow": 250000
3}
4

It makes Claude Code compact once the context reaches that size. It's capped at the model's own window, so on a 200k model it changes nothing.

I set it on the evening of September 25 and compared the week after (September 26 to October 2) with the 15 days before (September 10 to 24). The average context of my main sessions went from about 268k to 131k, and the cost of a request dropped by 39%. That's no coincidence: the context roughly halved, and every turn re-reads it. Cost fell a bit less than the context: output per request even went up. The daily total fell less, about a fifth, because I ran about twice as many subagent requests that week.

To be honest, 7 days is short, and the work in those days wasn't the same as in the weeks before. Subagent context fell too (153k to 92k), and I can't tell how much of that is the setting and how much is the work. Compaction also isn't free: it loses detail, and sometimes the agent reads a file again that it already had. I've now lowered it to 200000 and will measure again.

Compact-earlier advice is everywhere, at 50% or at 20%. I hadn't seen a before and after for it, so here's mine.

What my mods are actually for

Mods are TypeScript functions that run inside Claude Code. A shell hook runs before or after a call. A mod wraps it: on('tool.call', ($, e, next) => ...) can answer the call itself, retry it, wait for the result, start a new turn, or ask you something. A mod could save tokens too, for example by trimming big tool output before it lands in the context. Mine don't, they're about security and friction:

  • redact-secrets, security. It masks API keys and tokens in tool output before the model reads it, so a key the agent reads from a .env doesn't go to the model or stay in the transcript on disk. It knows the values in my env and .env* files plus common token formats. It doesn't touch what I type into the prompt, and a key in an unusual format can still get through. Zero tokens saved.
  • guard-autofix, less friction. It fixes calls my shell hooks would deny (a Co-Authored-By trailer, a long sleep, an em dash) instead of bouncing them back. About 75 retries in a month, roughly 0.1%.
  • ci-watch, workflow. It watches the pipeline after a push and wakes the agent once when it finishes. September had about 1,900 turns spent polling CI or sleeping. If every wait becomes a single wake-up, that saves 0.7%. Each poll is cheap, the session is warm and mostly cached.

I'd write all three again, just not for the cost. One more thing: mods run with your permissions and no sandbox, so read the code of any mod before you load it, mine included.

ci-watch: watches CI after a push and wakes the agent once
1// hooks/register.ts
2import type { EngineInterface, Register } from 'claude-code'
3import { type Forge, gitlabVerdict, githubVerdict, isRealPush, pushDirectory, type Verdict } from './ci'
4
5type Watch = { cwd: string; sha: string; branch: string; forge: Forge; polls: number }
6
7const POLL_MS = 60_000
8const GIVE_UP_WITHOUT_PIPELINE = 5
9const GIVE_UP_PENDING = 90
10
11async function git($: EngineInterface, cwd: string, args: string[]): Promise<string> {
12  const { exitCode, stdout } = await $.process.run(['git', ...args], { cwd })
13  if (exitCode !== 0) throw new Error(`git ${args[0]} failed`)
14  return stdout.trim()
15}
16
17async function describeHead($: EngineInterface, cwd: string): Promise<Watch> {
18  const [sha, branch, remote] = await Promise.all([
19    git($, cwd, ['rev-parse', 'HEAD']),
20    git($, cwd, ['rev-parse', '--abbrev-ref', 'HEAD']),
21    git($, cwd, ['remote', 'get-url', 'origin']),
22  ])
23  return { cwd, sha, branch, forge: remote.includes('github.com') ? 'github' : 'gitlab', polls: 0 }
24}
25
26async function verdictOf($: EngineInterface, w: Watch): Promise<Verdict> {
27  const argv =
28    w.forge === 'github'
29      ? ['gh', 'run', 'list', '--commit', w.sha, '--json', 'status,conclusion,url,name']
30      : ['glab', 'api', `projects/:fullpath/pipelines?sha=${w.sha}&per_page=1`]
31  const { stdout } = await $.process.run(argv, { cwd: w.cwd })
32  return w.forge === 'github' ? githubVerdict(stdout) : gitlabVerdict(stdout)
33}
34
35async function failedJobs($: EngineInterface, w: Watch, pipelineId: number): Promise<string> {
36  const { stdout } = await $.process.run(['glab', 'api', `projects/:fullpath/pipelines/${pipelineId}/jobs?scope[]=failed`], { cwd: w.cwd })
37  const jobs: unknown = JSON.parse(stdout)
38  if (!Array.isArray(jobs)) return ''
39  return jobs.map((job: { name?: string; web_url?: string }) => `${job.name ?? '?'} ${job.web_url ?? ''}`).join('\n')
40}
41
42async function report($: EngineInterface, w: Watch, verdict: Verdict & { state: 'passed' | 'failed' }): Promise<void> {
43  const label = `CI ${verdict.state} for ${w.branch} @ ${w.sha.slice(0, 8)}`
44  $.ui.status(label)
45  $.ui.toast(label)
46  const jobs = verdict.state === 'failed' && verdict.pipelineId !== undefined ? await failedJobs($, w, verdict.pipelineId).catch(() => '') : ''
47  const followUp = verdict.state === 'failed' ? 'Report it; fix only if fixing CI is part of the current task.' : 'No action needed unless you were waiting on it.'
48  await $.prompt.submit({ text: [`${label} (${verdict.detail}): ${verdict.url}`, jobs && `Failed jobs:\n${jobs}`, followUp].filter(Boolean).join('\n') })
49}
50
51async function tick($: EngineInterface, watches: Map<string, Watch>): Promise<void> {
52  for (const [sha, w] of watches) {
53    w.polls += 1
54    const verdict = await verdictOf($, w).catch((): Verdict => ({ state: 'pending' }))
55    const isStale = verdict.state === 'none' ? w.polls >= GIVE_UP_WITHOUT_PIPELINE : w.polls >= GIVE_UP_PENDING
56    if (verdict.state === 'passed' || verdict.state === 'failed') {
57      watches.delete(sha)
58      await report($, w, verdict)
59    } else if (isStale) {
60      watches.delete(sha)
61      $.ui.status(undefined)
62    } else {
63      $.ui.status(`CI ${verdict.state === 'none' ? 'waiting' : 'running'}: ${w.branch}`)
64    }
65  }
66}
67
68export const register: Register = on => {
69  const watches = new Map<string, Watch>()
70
71  on('session.start', async ($, e, next) => {
72    const started = await next(e)
73    $.clock.every(POLL_MS, () => void tick($, watches))
74    return started
75  })
76
77  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
78    const ran = await next(e)
79    if (ran.deny !== undefined || ran.isError === true || !isRealPush(e.command)) return ran
80    const w = await describeHead($, pushDirectory(e.command) ?? '.').catch(() => undefined)
81    if (w === undefined) return ran
82    watches.set(w.sha, w)
83    $.ui.status(`CI waiting: ${w.branch}`)
84    const note = `ci-watch is watching CI for ${w.branch} @ ${w.sha.slice(0, 8)} and will send you a message when it finishes. Do not poll the pipeline or wait on it.`
85    return { ...ran, context: [...(ran.context ?? []), note] }
86  })
87}
88
1// hooks/ci.ts
2export type Forge = 'gitlab' | 'github'
3
4export type Verdict =
5  | { state: 'none' | 'pending' }
6  | { state: 'passed' | 'failed'; url: string; detail: string; pipelineId?: number }
7
8const GITLAB_FINISHED = new Set(['success', 'failed', 'canceled', 'skipped', 'manual'])
9const GITHUB_OK = new Set(['success', 'skipped', 'neutral'])
10
11const isRecord = (value: unknown): value is Record<string, unknown> => typeof value === 'object' && value !== null
12
13const parseList = (json: string): Record<string, unknown>[] => {
14  try {
15    const parsed: unknown = JSON.parse(json)
16    return Array.isArray(parsed) ? parsed.filter(isRecord) : []
17  } catch {
18    return []
19  }
20}
21
22const text = (value: unknown): string => (typeof value === 'string' ? value : '')
23
24/** `glab api projects/:fullpath/pipelines?sha=..` output, newest first. */
25export const gitlabVerdict = (json: string): Verdict => {
26  const [latest] = parseList(json)
27  if (latest === undefined) return { state: 'none' }
28  const status = text(latest.status)
29  if (!GITLAB_FINISHED.has(status)) return { state: 'pending' }
30  const pipelineId = typeof latest.id === 'number' ? latest.id : undefined
31  return { state: status === 'success' ? 'passed' : 'failed', url: text(latest.web_url), detail: status, pipelineId }
32}
33
34/** `gh run list --commit .. --json status,conclusion,url,name` output: every workflow run of the commit. */
35export const githubVerdict = (json: string): Verdict => {
36  const runs = parseList(json)
37  if (runs.length === 0) return { state: 'none' }
38  if (runs.some(run => run.status !== 'completed')) return { state: 'pending' }
39  const failed = runs.filter(run => !GITHUB_OK.has(text(run.conclusion)))
40  if (failed.length === 0) return { state: 'passed', url: text(runs[0]?.url), detail: `${runs.length} runs` }
41  return { state: 'failed', url: text(failed[0]?.url), detail: failed.map(run => `${text(run.name)}: ${text(run.conclusion)}`).join(', ') }
42}
43
44/** The directory a push command runs git in: `git -C dir push`, else the last `cd dir` before it. */
45export const pushDirectory = (command: string): string | undefined => {
46  const push = /\bgit\b(?:\s+-C\s+(\S+))?\s+push\b/.exec(command)
47  if (push === null) return undefined
48  if (push[1] !== undefined) return push[1]
49  const cds = [...command.slice(0, push.index).matchAll(/(?:^|&&|;|\|\|)\s*cd\s+(["']?)([^\s"';&|]+)\1/g)]
50  return cds.at(-1)?.[2] ?? '.'
51}
52
53export const isRealPush = (command: string): boolean => pushDirectory(command) !== undefined && !/--dry-run\b/.test(command)
54
guard-autofix: fixes calls a hook would deny
1// hooks/register.ts
2import type { EngineInterface, Register } from 'claude-code'
3import { deepReplaceDashes, fixAuthoredCommand, isLongSleep, replaceDashes, routedSubagent } from './fixes'
4import { CURRENT, EVERY, isProtected, overwrittenBranches, pushDirectory } from './push'
5
6const MCP_WRITE = /^mcp__.*(send|draft|create|update|reply|forward|comment|add|batch)/
7
8async function checkedOutBranch($: EngineInterface, cwd: string | undefined): Promise<string> {
9  const { exitCode, stdout } = await $.process.run(['git', 'rev-parse', '--abbrev-ref', 'HEAD'], { cwd }).catch(() => ({ exitCode: 1, stdout: '' }))
10  // An unknown branch counts as protected
11  return exitCode === 0 ? stdout.trim() : EVERY
12}
13
14export const register: Register = on => {
15  // Pushes run freely; force-updating or deleting a shared branch never does, bypass mode included.
16  // a bare `git push -f` reads the branch from -C or the last cd, else the session's directory, not the shell's
17  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
18    const overwritten = overwrittenBranches(e.command)
19    if (overwritten.length === 0) return next(e)
20    const branches = await Promise.all(overwritten.map(branch => (branch === CURRENT ? checkedOutBranch($, pushDirectory(e.command)) : branch)))
21    const hit = branches.find(isProtected)
22    if (hit === undefined) return next(e)
23    return { deny: `Force-pushing or deleting ${hit === EVERY ? 'every branch' : hit} is blocked. Push a regular commit on top instead, or ask the user to run it themselves.` }
24  })
25
26  on('tool.call', { tool: 'Bash' }, ($, e, next) => {
27    const command = fixAuthoredCommand(e.command)
28    return isLongSleep(command) ? next({ ...e, command, run_in_background: true }) : next({ ...e, command })
29  })
30
31  on('tool.call', { tool: 'Edit' }, ($, e, next) => next({ ...e, new_string: replaceDashes(e.new_string) }))
32
33  on('tool.call', { tool: 'Write' }, ($, e, next) => next({ ...e, content: replaceDashes(e.content) }))
34
35  on('tool.call', ($, e, next) =>
36    MCP_WRITE.test(e.tool) ? next(deepReplaceDashes(e)) : next(e),
37  )
38
39  // The project's guard-subagent hook names the roster agent in its deny; take its advice instead of bouncing.
40  on('classic.PreToolUse', async ($, e, next) => {
41    const decision = await next(e)
42    if (e.tool !== 'Agent' || decision.deny === undefined) return decision
43    const agent = routedSubagent(decision.deny)
44    if (agent === undefined) return decision
45    const { tool, tool_use_id, ...args } = e
46    $.ui.toast(`guard-autofix: rerouted subagent to ${agent}`)
47    return {
48      updatedInput: { ...args, subagent_type: agent },
49      additionalContext: [`guard-subagent rerouted this Agent call to subagent_type "${agent}".`],
50    }
51  })
52}
53
1// hooks/fixes.ts
2const LONG_SLEEP = /(^|[^\w])sleep\s+([5-9]|[1-9]\d+)/
3const AUTHORED_TEXT = /\b(git\s+commit|gh\s+pr|glab\s+mr)\b/
4const ATTRIBUTION_LINE =
5  /^[ \t]*(co-authored-by:[^\n]*claude[^\n]*|claude-session:[^\n]*|https:\/\/claude\.ai\/code\/session_\S*|(🤖\s*)?generated with \[?claude code[^\n]*)\r?\n?/gim
6const ROUTE_HINT = /Use the `([\w-]+)` subagent for this task/
7
8export const isLongSleep = (command: string): boolean => LONG_SLEEP.test(command)
9
10export const replaceDashes = (text: string): string =>
11  text.replace(/[ \t]+[\u2014\u2013][ \t]+/g, ' - ').replace(/[\u2014\u2013]/g, '-')
12
13export const stripAttribution = (text: string): string =>
14  text.replace(ATTRIBUTION_LINE, '')
15
16export const fixAuthoredCommand = (command: string): string =>
17  AUTHORED_TEXT.test(command) ? replaceDashes(stripAttribution(command)) : command
18
19export const deepReplaceDashes = <T>(value: T): T => {
20  if (typeof value === 'string') return replaceDashes(value) as T
21  if (Array.isArray(value)) return value.map(deepReplaceDashes) as T
22  if (value !== null && typeof value === 'object')
23    return Object.fromEntries(Object.entries(value).map(([k, v]) => [k, deepReplaceDashes(v)])) as T
24  return value
25}
26
27export const routedSubagent = (denyReason: string): string | undefined =>
28  ROUTE_HINT.exec(denyReason)?.[1]
29
1// hooks/push.ts
2const PROTECTED = /^(main|master|dev|develop|development|stage|staging|prod|production)$/
3const PUSH = /\bgit\b(?:\s+-C\s+(\S+))?\s+push\b(.*)/
4const FORCE = /^(--force|--force-with-lease(=.*)?|--force-if-includes|-[a-z]*f[a-z]*)$/
5const DELETE = /^(--delete|-[a-z]*d[a-z]*)$/
6const TAKES_VALUE = /^(-o|--push-option|--repo|--receive-pack|--exec)$/
7
8/** Stands for the checked-out branch, which a push without a refspec updates */
9export const CURRENT = 'HEAD'
10/** Stands for every branch, as with --mirror or --all */
11export const EVERY = '*'
12
13export const isProtected = (branch: string): boolean =>
14  branch === EVERY || PROTECTED.test(branch.replace(/^refs\/heads\//, ''))
15
16/** Branches each `git push` in the command would force-update or delete */
17export function overwrittenBranches(command: string): string[] {
18  return command.split(/&&|\|\||;|\n|\|/).flatMap(segment => {
19    const push = PUSH.exec(segment)
20    if (push === null) return []
21    const words = (push[2] ?? '').trim().split(/\s+/).filter(Boolean).map(word => word.replace(/^["']|["']$/g, ''))
22    const flags = words.filter(word => word.startsWith('-'))
23    const positional = words.filter((word, i) => !word.startsWith('-') && !TAKES_VALUE.test(words[i - 1] ?? ''))
24    const force = flags.some(flag => FORCE.test(flag))
25    const remove = flags.some(flag => DELETE.test(flag))
26    if (flags.includes('--mirror') || (flags.includes('--all') && force)) return [EVERY]
27    const refspecs = positional.slice(1)
28    if (refspecs.length === 0) return force ? [CURRENT] : []
29    return refspecs.flatMap(refspec => {
30      const plus = refspec.startsWith('+')
31      const [source = '', target] = refspec.replace(/^\+/, '').split(':')
32      if (!force && !plus && !remove && source !== '') return []
33      return [target || source]
34    })
35  })
36}
37
38/** Where git runs: `git -C dir push`, else the last `cd dir` before it, else the session's directory */
39export function pushDirectory(command: string): string | undefined {
40  const push = PUSH.exec(command)
41  if (push === null) return undefined
42  if (push[1] !== undefined) return push[1]
43  const cds = [...command.slice(0, push.index).matchAll(/(?:^|&&|;|\|\|)\s*cd\s+(["']?)([^\s"';&|]+)\1/g)]
44  return cds.at(-1)?.[2]
45}
46
redact-secrets: masks secrets in tool output before the model reads it
1// hooks/register.ts
2import type { EngineInterface, Register } from 'claude-code'
3import { clipboardSecret, type KnownSecrets, parseSecrets, redactBlock, sourcedFiles } from './redact'
4
5const OWN_WORDS = new Set(['prompt', 'response'])
6
7async function learnFile($: EngineInterface, known: KnownSecrets, path: string): Promise<void> {
8  const home = (await $.env.get('HOME')) ?? ''
9  await $.fs
10    .read(path.replace(/^~(?=\/)/, home))
11    .then(text => parseSecrets(text, known))
12    .catch(() => undefined)
13}
14
15async function learnClipboard($: EngineInterface, known: KnownSecrets): Promise<void> {
16  const { stdout } = await $.process.run(['pbpaste']).catch(() => ({ stdout: '' }))
17  const secret = clipboardSecret(stdout)
18  if (secret !== undefined) known.set(secret, 'clipboard')
19}
20
21export const register: Register = on => {
22  const known: KnownSecrets = new Map()
23
24  on('session.start', async ($, e, next) => {
25    const started = await next(e)
26    // env files found once per session start, depth 4; sourced files are learned per call below.
27    await $.process
28      .run(['env'])
29      .then(({ stdout }) => parseSecrets(stdout, known))
30      .catch(() => undefined)
31    const found = await $.process
32      .run(['find', started.cwd, '-maxdepth', '4', '-name', 'node_modules', '-prune', '-o', '-name', '.env*', '-not', '-name', '*.example', '-type', 'f', '-print'])
33      .catch(() => ({ stdout: '' }))
34    await Promise.all(found.stdout.split('\n').filter(Boolean).map(path => learnFile($, known, path)))
35    return started
36  })
37
38  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
39    await Promise.all(sourcedFiles(e.command).map(path => learnFile($, known, path)))
40    if (/\bpbpaste\b/.test(e.command)) await learnClipboard($, known)
41    return next(e)
42  })
43
44  on('session.append', ($, e, next) =>
45    OWN_WORDS.has(e.door)
46      ? next(e)
47      : next({ ...e, message: { ...e.message, content: e.message.content.map(block => redactBlock(block, known)) } }),
48  )
49}
50
1// hooks/redact.ts
2import type { ApiContentBlock } from 'claude-code'
3
4export type KnownSecrets = Map<string, string>
5
6const SECRET_NAME = /(SECRET|TOKEN|PASS(WORD|WD)?|PRIVATE|CREDENTIAL|API_?KEY|ACCESS_?KEY|AUTH|SIGNING|SALT|COOKIE|WEBHOOK)/i
7const PUBLIC_NAME = /^(NEXT_PUBLIC_|VITE_|PUBLIC_)|_(PATH|FILE|ID|URL|HOST|PORT|ENABLED|TTL)$/i
8
9const PATTERNS: ReadonlyArray<[RegExp, string]> = [
10  [/-----BEGIN [A-Z ]*PRIVATE KEY-----[\s\S]*?-----END [A-Z ]*PRIVATE KEY-----/g, '[redacted:private-key]'],
11  [/\b(AKIA|ASIA)[0-9A-Z]{16}\b/g, '[redacted:aws-key-id]'],
12  [/\b(gh[pousr]_[A-Za-z0-9]{36,}|github_pat_\w{40,}|glpat-[\w-]{20,}|xox[abprs]-[\w-]{10,}|sk-(ant-)?[\w-]{20,})/g, '[redacted:token]'],
13  [/\beyJ[\w-]{8,}\.eyJ[\w-]{8,}\.[\w-]{8,}/g, '[redacted:jwt]'],
14  [/(\b[a-z][\w+.-]*:\/\/[^:/\s@]+:)[^@\s/]+@/gi, '$1[redacted]@'],
15  [/(authorization:\s*(bearer|basic)\s+)[^\s"']+/gi, '$1[redacted]'],
16  [/(aws_secret_access_key\s*[=:]\s*)\S+/gi, '$1[redacted]'],
17  [
18    /\b([A-Z0-9_]*(SECRET|TOKEN|PASSWORD|PASSWD|PRIVATE_KEY|API_KEY|ACCESS_KEY)[A-Z0-9_]*["']?\s*[=:]\s*["']?)[^\s"',;]{8,}/g,
19    '$1[redacted]',
20  ],
21]
22
23const unquote = (value: string): string => value.trim().replace(/^(['"])(.*)\1$/, '$2')
24
25const isSecretEntry = (name: string, value: string): boolean =>
26  SECRET_NAME.test(name) &&
27  !PUBLIC_NAME.test(name) &&
28  looksRandom(value)
29
30// A dictionary word like `postgres` would mask every mention of it; only mask values that look generated.
31const looksRandom = (value: string): boolean =>
32  value.length >= 16 || (value.length >= 8 && /\d/.test(value) && /[A-Za-z]/.test(value))
33
34/** `KEY=value` lines (an env file, `env` output) whose key names a secret. */
35export const parseSecrets = (text: string, into: KnownSecrets = new Map()): KnownSecrets => {
36  for (const line of text.split('\n')) {
37    const match = /^\s*(?:export\s+)?([A-Za-z_][A-Za-z0-9_]*)\s*=(.*)$/.exec(line)
38    const [, name, raw] = match ?? []
39    if (name === undefined || raw === undefined) continue
40    const value = unquote(raw)
41    if (isSecretEntry(name, value)) into.set(value, name)
42  }
43  return into
44}
45
46export const redact = (text: string, known: KnownSecrets): string => {
47  let out = text
48  for (const [value, name] of [...known].sort(([a], [b]) => b.length - a.length))
49    out = out.split(value).join(`[redacted:${name}]`)
50  for (const [pattern, replacement] of PATTERNS) out = out.replace(pattern, replacement)
51  return out
52}
53
54/** What `pbpaste` hands a command, when it is one token (a secret the user copied), not prose. */
55export const clipboardSecret = (text: string): string | undefined => {
56  const value = text.trim()
57  return value.length >= 8 && !/\s/.test(value) ? value : undefined
58}
59
60/** Files a shell command sources (`source f`, `. f`), the ones `set -a` loads secrets from. */
61export const sourcedFiles = (command: string): string[] =>
62  [...command.matchAll(/(?:^|[\s;&|(])(?:source|\.)\s+(["']?)([^\s;&|"')]+)\1/g)].flatMap(m => m[2] ?? [])
63
64const isTextBlock = (block: ApiContentBlock): block is ApiContentBlock & { text: string } =>
65  block.type === 'text' && typeof block.text === 'string'
66
67const redactText = (block: ApiContentBlock, known: KnownSecrets): ApiContentBlock =>
68  isTextBlock(block) ? { ...block, text: redact(block.text, known) } : block
69
70export const redactBlock = (block: ApiContentBlock, known: KnownSecrets): ApiContentBlock => {
71  if (block.type !== 'tool_result') return redactText(block, known)
72  const { content } = block
73  if (typeof content === 'string') return { ...block, content: redact(content, known) }
74  if (Array.isArray(content)) return { ...block, content: content.map((inner: ApiContentBlock) => redactText(inner, known)) }
75  return block
76}
77

The fourth mod: coming back after a break

The second thing that stood out was cache writes, 20% of the cost. On a Claude subscription, within the plan's included usage, Claude Code keeps the main session's prompt cache for an hour (otherwise it's 5 minutes by default), and writing to it costs 2x. Of all the big cache writes (over 30k tokens), 43% came right after a gap longer than an hour. That's lunch, a meeting, or the next morning: I come back to a big session, type one more prompt, and the whole context gets written to the cache again at double price, before the agent does anything.

The fix is something I already knew and never did: compact or /clear before going on. This is where a mod does help, not by saving tokens itself but by asking at the right moment. idle-compact checks every prompt I type. If the session was quiet for more than 55 minutes and holds more than 80k tokens, it asks whether to compact first. If I say yes, it drops my prompt, compacts with that prompt as the hint for what to keep, and sends the prompt again.

1import type { EngineInterface, Register } from 'claude-code'
2
3// The 1h prompt cache is gone after this much quiet; the next request re-writes the whole context at 2x.
4const CACHE_TTL_MS = 55 * 60_000
5const MIN_CONTEXT_TOKENS = 80_000
6const COMPACT = 'Compact first'
7const SEND = 'Send as is'
8const CANCEL = 'Cancel, new task (/clear)'
9const USER_ORIGINS = new Set(['composer', 'bridge'])
10
11export const register: Register = on => {
12  // In memory, so the first prompt after a restart or --resume isn't checked.
13  let lastActiveAt: number | undefined
14
15  on('turn.complete', async ($, e, next) => {
16    const done = await next(e)
17    if (e.agentId === undefined) lastActiveAt = await $.clock.now()
18    return done
19  })
20
21  on('prompt.submit', async ($, e, next) => {
22    const now = await $.clock.now()
23    const idleMs = lastActiveAt === undefined ? 0 : now - lastActiveAt
24    lastActiveAt = now
25    const isUsers = e.origin === undefined || USER_ORIGINS.has(e.origin.kind)
26    if (e.turnId !== undefined || !isUsers || e.text.trimStart().startsWith('/') || idleMs < CACHE_TTL_MS) return next(e)
27    const tokens = (await $.session.usage()).context.tokens ?? 0
28    if (tokens < MIN_CONTEXT_TOKENS) return next(e)
29    const answer = await $.ui
30      .ask(`Idle ${Math.round(idleMs / 60_000)} min, the prompt cache has expired. ${Math.round(tokens / 1000)}k tokens of context will be re-written at 2x. Compact first?`, {
31        header: 'idle-compact',
32        options: [COMPACT, SEND, CANCEL],
33      })
34      .catch(() => SEND)
35    if (answer === CANCEL) {
36      await $.prompt.fill({ text: e.text }).catch(() => undefined)
37      return { drop: 'idle-compact: prompt kept in the input box, run /clear first' }
38    }
39    if (answer !== COMPACT) return next(e)
40    // Compaction cannot run under the turn a prompt.submit hook holds: drop the prompt, compact, then send it again.
41    $.clock.after(0, () => void compactThenSend($, e.text))
42    return { drop: 'idle-compact: compacting first, your prompt follows' }
43  })
44}
45
46async function compactThenSend($: EngineInterface, text: string): Promise<void> {
47  const compacted = await $.session.compact({ instructions: `Keep what is needed for the next prompt: ${text.slice(0, 500)}` }).catch(() => undefined)
48  if (compacted === undefined || compacted.skip !== undefined) $.ui.toast('idle-compact: compaction did not run, sending as is')
49  await $.prompt.submit({ text })
50}
51

The tricky part is at the end. Claude Code refuses to compact from inside a prompt.submit hook, because the hook still holds the turn. So the mod drops the prompt, schedules the compaction for right after, and submits the prompt again itself.

September had about 150 of these returns, with an average context of 237k. If every one of them had been compacted first, I estimate it would have saved about 5% of the month. That's a rough estimate, not a measurement: it counts the compaction itself, but not the files the agent reads again afterwards. So the honest ranking is the window first, by far, and the idle compact second. The idea isn't new, there's a feature request for auto-compact on idle, and Zivtech wrote up the cache math. This is what it looked like in my logs.

Subagents are the next lever

Subagents went from about a quarter of the cost before the change to about half the week after, mostly because I ran twice as many of them. They run on the 5-minute cache, start cold, and some of mine still ran on Opus for work Sonnet does fine. So the next lever isn't a mod, it's picking the model per subagent: Sonnet for implementation from a plan, Haiku for simple lookups, Opus only where a mistake is expensive.

What it isn't

  • One person's month. My work is a lot of long sessions, browser automation and CI. Yours will look different, which is the reason to measure it yourself.
  • API prices, not my bill. I'm on a subscription. The weighting also ignores that Sonnet is cheaper than Opus, so a Sonnet-heavy setup gets overcounted.
  • Short after-period. 7 days after the setting change, different work than before. The direction is clear, the exact 39% isn't.
  • A rough estimate for idle-compact. It counts the compaction, not the files read again after it, and stops counting after 30 turns or 6 hours. It could be off in either direction.

Measure yours

My guess before measuring was that the agent's waiting was the expensive part. It was context size. Have you looked at where your agent's tokens go, and did it surprise you?