技术教程 · 阅读约 11 分钟

你的 AI Token 用量真的被浪费掉了——缓存命中率从 30% 提到 96% 的实战方法

你的 AI Token 用量真的被浪费掉了——缓存命中率从 30% 提到 96% 的实战方法

V2EX 上有人晒出了两个月 135 亿 Token 的用量,缓存命中率 96%。

很多人看到这个数字的第一反应是:这人真有钱。但我的第一反应是:他省了多少钱?

96% 的缓存命中意味着,135 亿次请求里,真正花钱的 Token 计算只有 5.4 亿次。剩下的 129.6 亿次,几乎是免费的。

而大多数人的项目,缓存命中率在 20%~40% 徘徊。原因不是模型不支持,是用法错了。


为什么你的 Prompt Cache 没生效

OpenAI 和 Anthropic 都支持 Prompt Caching,但触发条件有一个常见误区:

缓存是基于 prefix 的,不是整段 prompt。

你的 system prompt 必须是每次请求的前缀,且内容完全一致,缓存才会命中。很多人习惯把动态内容(时间戳、用户 ID、本次上下文)混进 system prompt,导致每次 prefix 都不同,缓存永远 miss。

还有一个坑:OpenAI 的缓存有最小长度限制,至少 1024 tokens 才会被缓存。你的 system prompt 如果只有两三句话,压根不会触发。


实战:构建一个缓存友好的 Code Review 工具

下面这个例子,把一个完整 repo 的代码作为静态 context 放进 system prompt,然后对它反复提问。整个过程只有第一次请求会真正计费 full token,后续全部命中缓存。


import os

import glob

from openai import OpenAI



# 用无量Api,国内直连,比官方便宜 60%+

client = OpenAI(

    api_key=os.environ.get("WULIANG_API_KEY"),

    base_url="https://api2everything.xyz/v1"

)



def load_repo_context(repo_path: str, extensions: list[str] = None) -> str:

    """把整个 repo 打包成 context string"""

    if extensions is None:

        extensions = [".py", ".ts", ".js", ".go", ".java"]



    files = []

    for ext in extensions:

        files.extend(glob.glob(f"{repo_path}/**/*{ext}", recursive=True))



    # 过滤掉依赖目录

    ignore = {"node_modules", ".git", "__pycache__", "dist", "build", ".venv"}

    files = [f for f in files if not any(d in f for d in ignore)]



    context_parts = []

    total_chars = 0



    for filepath in sorted(files):

        try:

            with open(filepath, "r", encoding="utf-8", errors="ignore") as f:

                content = f.read()

                # 单文件超过 50KB 截断,避免 context 爆炸

                if len(content) > 50_000:

                    content = content[:50_000] + "\n... [truncated]"

                rel_path = os.path.relpath(filepath, repo_path)

                context_parts.append(f"### File: {rel_path}\n```\n{content}\n```")

                total_chars += len(content)

        except Exception:

            continue



    print(f"Loaded {len(files)} files, ~{total_chars // 4} tokens estimated")

    return "\n\n".join(context_parts)





class RepoChatSession:

    """

    把 repo context 固定在 system prompt 开头,保证 prefix 不变。

    后续所有提问都复用同一份缓存。

    """



    def __init__(self, repo_path: str, model: str = "gpt-4.1"):

        self.model = model

        self.repo_context = load_repo_context(repo_path)



        # 静态 system prompt:repo 代码在前,指令在后

        # 注意:动态内容绝对不要放这里

        self.system_prompt = f"""You are a senior code reviewer with deep knowledge of software engineering best practices.



Below is the complete source code of the repository you are reviewing:



{self.repo_context}



---

Your responsibilities:

- Answer questions about this codebase accurately

- Identify bugs, security issues, and performance problems when asked

- Suggest concrete improvements with code examples

- Be direct and specific, not generic

"""

        self.conversation_history: list[dict] = []

        print(f"System prompt ready: ~{len(self.system_prompt) // 4} tokens")



    def ask(self, question: str) -> str:

        self.conversation_history.append({

            "role": "user",

            "content": question

        })



        response = client.chat.completions.create(

            model=self.model,

            messages=[

                {"role": "system", "content": self.system_prompt},

                *self.conversation_history

            ],

            temperature=0.3

        )



        answer = response.choices[0].message.content

        self.conversation_history.append({

            "role": "assistant",

            "content": answer

        })



        # 打印 token 使用情况,观察缓存命中

        usage = response.usage

        if hasattr(usage, "prompt_tokens_details"):

            cached = getattr(usage.prompt_tokens_details, "cached_tokens", 0)

            print(f"[Token Usage] prompt={usage.prompt_tokens}, "

                  f"cached={cached}, "

                  f"completion={usage.completion_tokens}, "

                  f"cache_hit_rate={cached/usage.prompt_tokens*100:.1f}%")



        return answer





# 使用示例

if __name__ == "__main__":

    session = RepoChatSession(repo_path="./your-project")



    questions = [

        "这个 repo 里有哪些潜在的 SQL 注入风险?",

        "找出所有没有做错误处理的异步函数",

        "auth 模块的 token 验证逻辑有什么问题?",

        "哪些函数的圈复杂度可能超过 10?",

    ]



    for q in questions:

        print(f"\n>>> {q}")

        answer = session.ask(q)

        print(answer)

        print("-" * 60)


关键设计点

1. System prompt 必须保持静态

repo 代码在 session 初始化时加载一次,之后不变。用户问题只出现在 messages 里,不污染 prefix。

2. 对话历史追加,不重置

conversation_history 不断 append,但每次请求时 system prompt 始终是同一个对象。这样 prefix 缓存在整个对话中持续有效。

3. 观察 cached_tokens 字段

第一次请求会看到 cached=0,从第二次开始应该看到大量 cached tokens。如果一直是 0,说明 prefix 有变动,需要排查。

4. Token 估算

平均一个中文字符约 1.5~2 token,英文代码约 3~4 字符/token。一个中型 Python 项目(200 个文件)大概在 80k~150k tokens 之间,完全在 GPT-4.1 的 context 窗口内。


实际省钱幅度

以 GPT-4.1 为例,官方定价:

  • 普通 input: $2 / 1M tokens
  • cached input: $0.50 / 1M tokens(75% 折扣)

100 次问答,每次 100k token prefix,命中率 96%:

  • 无缓存:$20
  • 有缓存:$0.8 × 4% full price + $0.50 × 96% cache = 约 $1.6

省了 92%。这就是那个 135 亿 token 账单背后的逻辑。


关于 API 费用

国内访问 OpenAI 官方 API 还有一个隐性成本:代理稳定性。断连、超时、重试,这些都在消耗你的实际 token(重试的请求不退费)。

我现在用的是 无量Api,国内直连,支持 OpenAI / Claude / Gemini / DeepSeek 等 300+ 模型,价格比官方便宜约 65%,改一行 base_url 就能切过去,注册还送 ¥1 余额直接测试。


把这个脚本跑起来,第二个问题开始观察 cache_hit_rate,如果没到 90% 以上,欢迎评论区说说你的 system prompt 结构,一起排查。

觉得有用的话点个赞,后续可以写「如何用 Anthropic 的 explicit cache breakpoints 做更精细的缓存控制」。