cache-smith
LLM caching optimizer — analyze KV prompt caching potential and simulate semantic caching savings before committing to a full gateway.
A CLI tool that helps developers optimize LLM API costs through caching.
Three modes:
-
cache-smith analyze— analyzes prompt templates for provider-side KV prompt caching potential. Measures prefix overlap, scores cacheability, and estimates cost savings across OpenAI, Anthropic, and Google. -
cache-smith simulate— simulates client-side semantic caching using local embeddings (sentence-transformers). Runs similarity sweeps, finds optimal thresholds, and reports hit rates and savings. -
cache-smith proxy— lightweight HTTP proxy that intercepts OpenAI-compatible API calls and serves cached responses via semantic matching. Works as a drop-in replacement for your existing base URL.
Key features:
- Local embeddings (no API calls needed) with deterministic hash fallback
- Provider-aware cost modeling (OpenAI, Anthropic, Google)
- Optimal threshold suggestion via sweep (0.70-0.98)
- Rich terminal tables + JSON output
- Thread-safe in-memory proxy cache with
X-Cacheheaders
Install: pip install -e . from source