你的 AI Token 用量真的被浪费掉了——缓存命中率从 30% 提到 96% 的实战方法
你的 AI Token 用量真的被浪费掉了——缓存命中率从 30% 提到 96% 的实战方法
V2EX 上有人晒出了两个月 135 亿 Token 的用量,缓存命中率 96%。
很多人看到这个数字的第一反应是:这人真有钱。但我的第一反应是:他省了多少钱?
96% 的缓存命中意味着,135 亿次请求里,真正花钱的 Token 计算只有 5.4 亿次。剩下的 129.6 亿次,几乎是免费的。
而大多数人的项目,缓存命中率在 20%~40% 徘徊。原因不是模型不支持,是用法错了。
为什么你的 Prompt Cache 没生效
OpenAI 和 Anthropic 都支持 Prompt Caching,但触发条件有一个常见误区:
缓存是基于 prefix 的,不是整段 prompt。
你的 system prompt 必须是每次请求的前缀,且内容完全一致,缓存才会命中。很多人习惯把动态内容(时间戳、用户 ID、本次上下文)混进 system prompt,导致每次 prefix 都不同,缓存永远 miss。
还有一个坑:OpenAI 的缓存有最小长度限制,至少 1024 tokens 才会被缓存。你的 system prompt 如果只有两三句话,压根不会触发。
实战:构建一个缓存友好的 Code Review 工具
下面这个例子,把一个完整 repo 的代码作为静态 context 放进 system prompt,然后对它反复提问。整个过程只有第一次请求会真正计费 full token,后续全部命中缓存。
import os
import glob
from openai import OpenAI
# 用无量Api,国内直连,比官方便宜 60%+
client = OpenAI(
api_key=os.environ.get("WULIANG_API_KEY"),
base_url="https://api2everything.xyz/v1"
)
def load_repo_context(repo_path: str, extensions: list[str] = None) -> str:
"""把整个 repo 打包成 context string"""
if extensions is None:
extensions = [".py", ".ts", ".js", ".go", ".java"]
files = []
for ext in extensions:
files.extend(glob.glob(f"{repo_path}/**/*{ext}", recursive=True))
# 过滤掉依赖目录
ignore = {"node_modules", ".git", "__pycache__", "dist", "build", ".venv"}
files = [f for f in files if not any(d in f for d in ignore)]
context_parts = []
total_chars = 0
for filepath in sorted(files):
try:
with open(filepath, "r", encoding="utf-8", errors="ignore") as f:
content = f.read()
# 单文件超过 50KB 截断,避免 context 爆炸
if len(content) > 50_000:
content = content[:50_000] + "\n... [truncated]"
rel_path = os.path.relpath(filepath, repo_path)
context_parts.append(f"### File: {rel_path}\n```\n{content}\n```")
total_chars += len(content)
except Exception:
continue
print(f"Loaded {len(files)} files, ~{total_chars // 4} tokens estimated")
return "\n\n".join(context_parts)
class RepoChatSession:
"""
把 repo context 固定在 system prompt 开头,保证 prefix 不变。
后续所有提问都复用同一份缓存。
"""
def __init__(self, repo_path: str, model: str = "gpt-4.1"):
self.model = model
self.repo_context = load_repo_context(repo_path)
# 静态 system prompt:repo 代码在前,指令在后
# 注意:动态内容绝对不要放这里
self.system_prompt = f"""You are a senior code reviewer with deep knowledge of software engineering best practices.
Below is the complete source code of the repository you are reviewing:
{self.repo_context}
---
Your responsibilities:
- Answer questions about this codebase accurately
- Identify bugs, security issues, and performance problems when asked
- Suggest concrete improvements with code examples
- Be direct and specific, not generic
"""
self.conversation_history: list[dict] = []
print(f"System prompt ready: ~{len(self.system_prompt) // 4} tokens")
def ask(self, question: str) -> str:
self.conversation_history.append({
"role": "user",
"content": question
})
response = client.chat.completions.create(
model=self.model,
messages=[
{"role": "system", "content": self.system_prompt},
*self.conversation_history
],
temperature=0.3
)
answer = response.choices[0].message.content
self.conversation_history.append({
"role": "assistant",
"content": answer
})
# 打印 token 使用情况,观察缓存命中
usage = response.usage
if hasattr(usage, "prompt_tokens_details"):
cached = getattr(usage.prompt_tokens_details, "cached_tokens", 0)
print(f"[Token Usage] prompt={usage.prompt_tokens}, "
f"cached={cached}, "
f"completion={usage.completion_tokens}, "
f"cache_hit_rate={cached/usage.prompt_tokens*100:.1f}%")
return answer
# 使用示例
if __name__ == "__main__":
session = RepoChatSession(repo_path="./your-project")
questions = [
"这个 repo 里有哪些潜在的 SQL 注入风险?",
"找出所有没有做错误处理的异步函数",
"auth 模块的 token 验证逻辑有什么问题?",
"哪些函数的圈复杂度可能超过 10?",
]
for q in questions:
print(f"\n>>> {q}")
answer = session.ask(q)
print(answer)
print("-" * 60)
关键设计点
1. System prompt 必须保持静态
repo 代码在 session 初始化时加载一次,之后不变。用户问题只出现在 messages 里,不污染 prefix。
2. 对话历史追加,不重置
conversation_history 不断 append,但每次请求时 system prompt 始终是同一个对象。这样 prefix 缓存在整个对话中持续有效。
3. 观察 cached_tokens 字段
第一次请求会看到 cached=0,从第二次开始应该看到大量 cached tokens。如果一直是 0,说明 prefix 有变动,需要排查。
4. Token 估算
平均一个中文字符约 1.5~2 token,英文代码约 3~4 字符/token。一个中型 Python 项目(200 个文件)大概在 80k~150k tokens 之间,完全在 GPT-4.1 的 context 窗口内。
实际省钱幅度
以 GPT-4.1 为例,官方定价:
- 普通 input: $2 / 1M tokens
- cached input: $0.50 / 1M tokens(75% 折扣)
100 次问答,每次 100k token prefix,命中率 96%:
- 无缓存:$20
- 有缓存:$0.8 × 4% full price + $0.50 × 96% cache = 约 $1.6
省了 92%。这就是那个 135 亿 token 账单背后的逻辑。
关于 API 费用
国内访问 OpenAI 官方 API 还有一个隐性成本:代理稳定性。断连、超时、重试,这些都在消耗你的实际 token(重试的请求不退费)。
我现在用的是 无量Api,国内直连,支持 OpenAI / Claude / Gemini / DeepSeek 等 300+ 模型,价格比官方便宜约 65%,改一行 base_url 就能切过去,注册还送 ¥1 余额直接测试。
把这个脚本跑起来,第二个问题开始观察 cache_hit_rate,如果没到 90% 以上,欢迎评论区说说你的 system prompt 结构,一起排查。
觉得有用的话点个赞,后续可以写「如何用 Anthropic 的 explicit cache breakpoints 做更精细的缓存控制」。