Persistent KV-cache for LLM coding agents
Problem
Every time a developer boots an LLM coding agent (e.g. opencode) into a codebase, the agent burns 40+ tool calls re-exploring the project to rebuild its context, because KV-cache state only lives during the work session and is discarded on exit. Re-establishing that context on every fresh start wastes significant time and tokens (and money) for anyone working on the same codebase day after day.
Opportunity
A tool or middleware that snapshots, stores and restores per-repository agent context (KV-state or an equivalent summarized exploration cache) so a fresh boot starts warm instead of re-exploring from zero, dramatically cutting latency and token cost for daily agent users.
Market analysis
The cold-start cost is real, but literal KV-cache snapshots only work for self-hosted models; for API users the win is a summarized exploration cache injected at boot. Native session-resume and memory plugins already attack this, so the wedge is cross-agent, repo-scoped warm context rather than cache persistence itself.
Market · Daily users of CLI coding agents (opencode, Claude Code, Codex); dev-tools segment with strong bottom-up adoption and high token-cost sensitivity.
Pricing · Comparables are free open-source plugins; monetization would likely be a hosted/team tier in the $10-30 per user per month range rather than direct sales.
Pros
- + Real, quantified pain: re-exploration burns tokens and latency on every boot.
- + Low distribution friction as a local OSS CLI or agent plugin.
- + Prompt caching economics are proven, with large reported cost cuts on cached prefixes.
Cons
- − Literal KV-cache persistence only works with local/self-hosted models; API users cannot access KV state.
- − Provider-side prompt caching plus agents' built-in session resume shrink the addressable gap.
- − Crowded OSS memory-plugin space (claude-mem, agentmemory) moving fast.
Existing / similar tools
- → claude-mem
- → grov ↗
- → agentmemory ↗
- → OpenCode (built-in session resume)
Source
Hacker News (Ask HN)
The hard truth is that KV-cache state is the wrong layer for most of this market: if your agent runs against Anthropic or OpenAI APIs, you never see the tensors — the provider’s prompt cache is the only cache you get, and Claude Code reportedly already achieves ~92% cache-hit rates on long agentic contexts. The defensible version of this idea is a repo-scoped exploration cache that works across agents and providers: capture what the agent learns about a codebase, compress it, and inject it on boot — which is exactly the race grov and claude-mem are already running. A solo builder’s edge would be doing this at the proxy level (intercepting agent traffic once, serving every agent) rather than shipping yet another per-agent plugin, with measurable per-repo token savings as the headline metric. Without that angle, this is a crowded OSS niche with weak willingness to pay.