UP TO 90% LOWER CACHED INPUT COST

Find the token
that kills cache.

Whether you trace through LangSmith or run agents directly in your IDE — CACHECATCH turns runs and sessions into a cache-loss report: exact divergence, wasted spend, and the prompt fix your team should ship first.

LangSmithLangfuseBraintrustLocal AgentsClaude CodeCodexOpenCode
$npx --yes cachecatch audit local --window 7d
Free + open-source . Runs locally. No prompts uploaded.
View sample report
Real CLI report

A cache audit
engineers can act on.

A fast CLI report that summarizes a sample audit — recoverable loss, the top leaking routes, the exact prefix break, and the prompt layout to ship next. Works with cloud traces and local IDE agent sessions alike.

╭ Cachecatch Local ────────────────────────────────────────────────────────────────────────────────────────╮ │ │ │ Cachecatch Local Agent Report │ │ Generated 2026-07-01T12:00:00Z · time window 7d │ │ │ │ Main finding: Cachecatch analyzed 89 local agent sessions across Claude Code, OpenCode, Codex. Token │ │ activity is mixed; cost telemetry is unavailable and pricing coverage is unavailable. The observed │ │ Claude Code, OpenCode cache profile is 9% (LOW). Missing cache fields are not counted as 0%. The main │ │ fixable signal is context hygiene in the transcript-visible sessions: git diffs, terminal output, and │ │ timestamps appear before repo instructions in 68% of sessions │ │ Cachecatch found 342 stored sessions total and 89 sessions in this time window. "Analyzed" means │ │ Cachecatch could parse the session format safely. │ │ │ │ ╭ Builder Scorecard ───────────────────────────────────────────────────────────────────────────────────╮ │ │ │ AGENTIC SESSIONS 89 parsed in this time window │ │ │ │ TOKEN ACTIVITY 2,840,000 mixed observed/estimated │ │ │ │ TOOL CALLS 1,247 observed where local storage exposes tool records │ │ │ │ SUBAGENT RUNS 38 explore (20), general (10), writer (6), editor (2) │ │ │ │ IDE AGENTS USED Claude Code, OpenCode, Codex │ │ │ │ MODELS DETECTED 4 │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ Cache Profile ───────────────────────────────────────────────────────────────────────────────────────╮ │ │ │ CACHE READ PROFILE 9% LOW │ │ │ │ CACHE READ TOKENS 186,000 observed │ │ │ │ CACHE WRITE TOKENS 412,000 observed │ │ │ │ CACHE-TOKEN TELEMETRY unavailable │ │ │ │ COST TELEMETRY unavailable │ │ │ │ PRICING COVERAGE unavailable │ │ │ │ TRANSCRIPT HYGIENE SIGNALS available │ │ │ │ │ │ │ │ Scope observed on Claude Code, OpenCode only; transcript-only agents are excluded. │ │ │ │ Trust CacheCatch does not treat missing cache telemetry as zero. Missing fields are reported as not │ │ │ │ reported. │ │ │ │ Limit Claude/Codex local transcripts are not counted as 0%. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ Context Hygiene ─────────────────────────────────────────────────────────────────────────────────────╮ │ │ │ CONTEXT HYGIENE SCORE 52 / 100 WATCH ████████████░░░░░░░░░░░░ │ │ │ │ EFFICIENCY UPSIDE SIGNAL HIGH based on evidence-backed hygiene findings │ │ │ │ RESEARCH BENCHMARK RANGE 41-80% cost reduction in agentic workloads; not this audit's claim │ │ │ │ KNOWN MODEL COST $12.40 equivalent observed where local telemetry exposes cost │ │ │ │ RECOVERABLE TOKEN GAP 12-18% signal dollar estimate unavailable because cost/pricing coverage is │ │ │ │ partial │ │ │ │ BILLING CONFIDENCE PARTIAL subscriptions/promos/enterprise pricing not visible │ │ │ │ │ │ │ │ Caveat Dollar values are model-cost equivalents only. │ │ │ │ Basis Unpriced or partially priced token opportunity is not converted to dollars. │ │ │ │ Pricing Conservative cache-discount range used where exact cached-input price was unavailable. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ Report Confidence ───────────────────────────────────────────────────────────────────────────────────╮ │ │ │ Confidence MEDIUM │ │ │ │ Medium confidence: token/cache direction is useful, but some agents or models are estimated or not │ │ │ │ priced. │ │ │ │ Use sessions, observed token/cache rows, model mix, and visible context-hygiene findings as the │ │ │ │ reliable parts. Unsafe dollar values are suppressed. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ■ IDE AGENT BREAKDOWN │ │ Stored = all local sessions found. In time window = sessions whose modified/update time is inside the │ │ requested time window. Analyzed = parsed safely. │ │ │ │ ╭ Claude Code (detected) ──────────────────────────────────────────────────────────────────────────────╮ │ │ │ Visibility telemetry-rich cache telemetry not reported │ │ │ │ Sessions 42 analyzed 42 in time window · 156 stored │ │ │ │ Tokens 1,420,000 observed from Claude local token events │ │ │ │ Cache telemetry not reported no explicit cache fields observed │ │ │ │ Cache read not reported missing cache field │ │ │ │ Input tokens 1,060,000 observed │ │ │ │ Cached input not reported missing cache field │ │ │ │ Output tokens 360,000 observed │ │ │ │ Cost $6.80 observed/estimated from telemetry │ │ │ │ Source OTel logs │ │ │ │ Context hygiene 48 / 100 WATCH │ │ │ │ Tool calls 624 │ │ │ │ Subagents 18 explore (12), general (6) │ │ │ │ Top models claude-sonnet-4-20250514, claude-haiku-3.5 ranked by parsed sessions │ │ │ │ Finding Git diffs and terminal output load before AGENTS.md rules in 71% of sessions, breaking │ │ │ │ prefix stability. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ OpenCode (detected) ─────────────────────────────────────────────────────────────────────────────────╮ │ │ │ Visibility telemetry-rich cache telemetry not reported │ │ │ │ Sessions 38 analyzed 38 in time window · 128 stored │ │ │ │ Tokens 980,000 observed from OpenCode local database │ │ │ │ Cache telemetry not reported no explicit cache fields observed │ │ │ │ Cache read not reported missing cache field │ │ │ │ Input tokens 740,000 observed │ │ │ │ Cached input not reported missing cache field │ │ │ │ Output tokens 240,000 observed │ │ │ │ Cost $4.10 observed/estimated from telemetry │ │ │ │ Source local database │ │ │ │ Context hygiene 58 / 100 WATCH │ │ │ │ Tool calls 486 │ │ │ │ Subagents 14 explore (8), writer (6) │ │ │ │ Top models gpt-4.1, gpt-4.1-mini ranked by parsed sessions │ │ │ │ Finding Timestamps and session metadata inject before stable project instructions, reducing prefix │ │ │ │ reuse. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ Codex (detected) ────────────────────────────────────────────────────────────────────────────────────╮ │ │ │ Visibility transcript-rich / cache-telemetry limited cache telemetry not reported │ │ │ │ Sessions 9 analyzed 9 in time window · 58 stored │ │ │ │ Tokens 440,000 estimated from transcript text length │ │ │ │ Cache telemetry not reported no explicit cache fields observed │ │ │ │ Cache read not reported missing cache field │ │ │ │ Input tokens 320,000 estimated │ │ │ │ Cached input not reported missing cache field │ │ │ │ Output tokens 120,000 estimated │ │ │ │ Cost $1.50 observed/estimated from telemetry │ │ │ │ Source transcript │ │ │ │ Context hygiene 64 / 100 WATCH │ │ │ │ Tool calls 137 │ │ │ │ Subagents 6 general (4) │ │ │ │ Top models codex-mini-latest ranked by parsed sessions │ │ │ │ Finding Cache telemetry not exposed in local Codex session files; hygiene score based on transcript │ │ │ │ analysis. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ■ TELEMETRY SETUP │ │ ╭ Enable Future Exact Telemetry ───────────────────────────────────────────────────────────────────────╮ │ │ │ Run `npx --yes cachecatch init claude`, then start Claude Code with the generated env file to enable │ │ │ │ future cache/token telemetry. │ │ │ │ Run `npx --yes cachecatch init codex` to enable future Codex OTel telemetry. │ │ │ │ Then run `npx --yes cachecatch daemon` while using the agent. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ■ TOP PROJECTS │ │ Project advice is based on the local project folders found in agent storage, plus AGENTS.md / CLAUDE.md │ │ presence. │ │ │ │ ╭ ~/code/support-copilot ──────────────────────────────────────────────────────────────────────────────╮ │ │ │ Sessions 38 │ │ │ │ Token activity 1,180,000 │ │ │ │ Cache read 9% observed │ │ │ │ AGENTS.md present keep stable │ │ │ │ CLAUDE.md present Claude-specific rules available │ │ │ │ → AGENTS.md is present and looks stable, but observed cache-read is only 9% — the leakage is below │ │ │ │ the instruction files, in task context. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ ~/code/docs-rag ─────────────────────────────────────────────────────────────────────────────────────╮ │ │ │ Sessions 28 │ │ │ │ Token activity 820,000 │ │ │ │ Cache read 7% observed │ │ │ │ AGENTS.md missing add stable repo rules │ │ │ │ CLAUDE.md missing optional unless Claude Code is used here │ │ │ │ → 28 sessions in this project started from a clean room; they re-built the same project context │ │ │ │ every run. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ ~/code/refund-service ───────────────────────────────────────────────────────────────────────────────╮ │ │ │ Sessions 23 │ │ │ │ Token activity 840,000 │ │ │ │ Cache read not reported not reported │ │ │ │ AGENTS.md present keep stable │ │ │ │ CLAUDE.md missing optional unless Claude Code is used here │ │ │ │ → Keep AGENTS.md byte-stable across sessions. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ■ TOP FIXABLE ISSUES │ │ │ │ ╭ Dynamic context loads before stable rules ───────────────────────────────────────────────────────────╮ │ │ │ What we found: git diffs, terminal output, and timestamps appear before repo instructions in 68% of │ │ │ │ sessions │ │ │ │ Why it matters: Prompt caching is prefix-sensitive. If diffs, logs, timestamps, and tool output │ │ │ │ appear before stable repo rules, the reusable prefix changes sooner. │ │ │ │ Where: Claude Code + OpenCode + Codex · 89 sessions · top projects: ~/code/support-copilot, │ │ │ │ ~/code/docs-rag │ │ │ │ What to do next: Put AGENTS.md/CLAUDE.md first, push diffs and logs to the tail. │ │ │ │ Validation: Rerun the same project/window and compare cache-read %, token activity basis, and │ │ │ │ whether this finding still appears. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ AGENTS.md missing or minimal in key projects ────────────────────────────────────────────────────────╮ │ │ │ What we found: docs-rag has no AGENTS.md; refund-service has no CLAUDE.md │ │ │ │ Why it matters: Agent instruction files are the best place for stable context: repo layout, │ │ │ │ commands, constraints, style, and testing expectations. │ │ │ │ Where: Claude Code + OpenCode + Codex · 89 sessions · top projects: ~/code/support-copilot, │ │ │ │ ~/code/docs-rag │ │ │ │ What to do next: Add stable repo rules to both projects. │ │ │ │ Validation: Rerun the same project/window and compare cache-read %, token activity basis, and │ │ │ │ whether this finding still appears. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ Repeated boilerplate across sessions ────────────────────────────────────────────────────────────────╮ │ │ │ What we found: Repo layout and command hints repeated in 19 of 89 sessions │ │ │ │ Why it matters: Repeating repo instructions inside ad hoc prompts wastes context and makes every │ │ │ │ agent rebuild the same prefix instead of reusing one stable project file. │ │ │ │ Where: Claude Code + OpenCode + Codex · 89 sessions · top projects: ~/code/support-copilot, │ │ │ │ ~/code/docs-rag │ │ │ │ What to do next: Use a shared instruction file instead of repeating boilerplate. │ │ │ │ Validation: Rerun the same project/window and compare cache-read %, token activity basis, and │ │ │ │ whether this finding still appears. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ Cache telemetry unavailable for Codex ───────────────────────────────────────────────────────────────╮ │ │ │ What we found: Codex local sessions do not expose cache-read/cache-write token fields │ │ │ │ Why it matters: This is a measurement limitation, not proof of bad Codex/Claude behavior. The honest │ │ │ │ output is 'not reported' until local files expose cache-read/cache-write tokens. │ │ │ │ Where: Claude Code + OpenCode + Codex · 89 sessions · top projects: ~/code/support-copilot, │ │ │ │ ~/code/docs-rag │ │ │ │ What to do next: Use transcript-based hygiene findings as the reliable signal. │ │ │ │ Validation: Rerun the same project/window and compare cache-read %, token activity basis, and │ │ │ │ whether this finding still appears. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ■ BEFORE / FIX / AFTER │ │ ╭ 1 BEFORE - Local Context Today ──────────────────────────────────────────────────────────────────────╮ │ │ │ × git diffs, terminal output, and timestamps appear before repo instructions in 68% of sessions │ │ │ │ × Prompt caching is prefix-sensitive. If diffs, logs, timestamps, and tool output appear before │ │ │ │ stable repo rules, the reusable prefix changes sooner. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ 2 CACHECATCH FIX ────────────────────────────────────────────────────────────────────────────────────╮ │ │ │ → Put stable repo identity, rules, tools, and output conventions in AGENTS.md / CLAUDE.md. │ │ │ │ → Keep that stable block byte-stable across sessions. │ │ │ │ → Put AGENTS.md/CLAUDE.md first, push diffs and logs to the tail. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ╭ 3 AFTER - Cache-Ready Local Agent Workflow ──────────────────────────────────────────────────────────╮ │ │ │ ✔ Stable prefix first: role, repo rules, architecture constraints, command policy. │ │ │ │ ✔ Dynamic tail last: task notes, terminal output, diffs, errors, current state. │ │ │ │ ✔ Validation: rerun audit and compare observed cache-read %, token basis, and whether the same │ │ │ │ finding remains. │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ■ PUBLIC SHARE SUMMARY │ │ ╭ Public Share Summary ────────────────────────────────────────────────────────────────────────────────╮ │ │ │ 89 agentic sessions in 7d │ │ │ │ 2.84M token activity analyzed · mixed observed/estimated │ │ │ │ 1,247 tool calls │ │ │ │ 38 subagent runs │ │ │ │ 9% observed cache-read profile │ │ │ │ 4 models detected │ │ │ │ │ │ │ │ npx --yes cachecatch share ./reports/<report>.json │ │ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────╯ │ │ │ │ ■ AGENT REPAIR PROMPT │ │ You are improving local agent prompt-cache hygiene. │ │ │ │ 1. Dynamic context loads before stable rules. │ │ Evidence: git diffs, terminal output, and timestamps appear before repo instructions in 68% of sessions │ │ Fix: Put AGENTS.md/CLAUDE.md first, push diffs and logs to the tail. │ │ │ │ 2. AGENTS.md missing or minimal in key projects. │ │ Evidence: docs-rag has no AGENTS.md; refund-service has no CLAUDE.md │ │ Fix: Add stable repo rules to both projects. │ │ │ │ Validate by rerunning CacheCatch for the same window and checking whether the same findings remain. │ │ │ │ ■ DISCLAIMER │ │ Dollar values are model-cost equivalents only. Subscriptions, promotions, and enterprise pricing are not │ │ visible. Cache-discount ranges are conservative where exact cached-input pricing was unavailable. │ │ │ ╰────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭ ♥ Support + Share ────────────────────────────────────────────────────────────────────────────────────────╮ │ Support Cachecatch: copy and run npx --yes cachecatch share to generate your share banner. │ │ It will ask for your X handle, make the PNG, and print ready-to-use X copy/link. │ ╰────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
Summarized from sample report. Run locally. No prompts / keys stored.
Why this matters

Provider caching is real.
Your prompt order can still break it.

OpenAI and Anthropic reward stable prefixes — whether you call them from cloud traces or local IDE agents. CACHECATCH shows exactly where your prompt stops being cacheable and what that costs.

up to 90%lower cached input-token cost documented by OpenAI.
up to 80%lower latency possible when reusable prefixes hit cache.
10%of standard input price for Anthropic cache-read tokens.
45-80%API cost reduction measured in a 2026 agentic prompt-caching evaluation.

Sources: OpenAI prompt caching docs, Anthropic pricing docs, and the 2026 agentic prompt-caching evaluation.

What the report gives you

The missing layer
after tracing.

Tracing tools show runs, latency, token usage, and cost. Local agents generate sessions but leave cache efficiency unseen. CACHECATCH turns both into the cache diagnosis your team can act on immediately.

1Find the cache breaker

Request IDs, timestamps, user metadata, RAG blocks, tool schemas, and dynamic system prompts that appear before the stable prefix.

2Rank routes by waste

Group repeated agent routes and show monthly waste, divergence depth, severity, evidence, and confidence per route.

3Ship the fix plan

Move stable rules, tools, policy, and examples first. Push session metadata, user query, and tool outputs into the dynamic tail.

Your tracing stack keeps

  • traces, runs, and latency
  • token usage and model metadata
  • debugging context and observability

CACHECATCH adds

  • first divergence token
  • cache-specific waste estimate
  • exact prompt-layout fix plan
Run the audit

Stop paying full price
for reusable context.

Grab the CLI command and run the audit on your local agent sessions or cloud traces in minutes.

$npx --yes cachecatch audit local --window 7d
Free + open-source . Runs locally. No prompts uploaded.