月花 200 刀调 Claude 的我,改了一行代码后现在只花 60
月花 200 刀调 Claude 的我,改了一行代码后现在只花 60
结论先说:把 base_url 换掉,账单直接降到三分之一。
痛点
我有个跑了半年的 code review bot,后端全是 Claude Sonnet。上个月账单 $217,看着有点难绷。
细拆了一下,问题很集中:
1. 所有请求都打 Sonnet,不管是"格式化这段代码"还是"解释这个架构",统一顶配
2. 没做任何缓存,同一段 system prompt 每次都重新计费
3. 直连 api.anthropic.com,国内偶尔超时重试,token 白费
其实第 3 条是最亏的——超时之后客户端自动重试,等于同一个请求付了两次钱,但我只拿到一次结果。
怎么改
第一步:换 base_url,解决超时重试问题
原来的代码:
import anthropic
client = anthropic.Anthropic(api_key="sk-ant-xxx")
改成:
import anthropic
client = anthropic.Anthropic(
api_key="你的中转 key",
base_url="https://api.api2everything.xyz/v1",
)
接口格式完全兼容,其他代码一行不动。用 OpenAI SDK 也一样,改 base_url 就完事。
第二步:按任务分流模型
不是所有任务都需要 Sonnet。我按复杂度拆成两档:
def get_model(task_type: str) -> str:
# 简单任务:格式化、注释、重命名
if task_type in ("format", "comment", "rename"):
return "claude-haiku-3-5" # 比 Sonnet 便宜 ~20x
# 复杂任务:架构分析、bug 定位
return "claude-sonnet-4-5"
def review(code: str, task_type: str = "general"):
resp = client.messages.create(
model=get_model(task_type),
max_tokens=1024,
messages=[{"role": "user", "content": code}],
)
return resp.content[0].text
第三步:system prompt 走 cache(Anthropic 原生支持)
长 system prompt 每次重新传很费钱,加一个 cache_control 就能复用:
response = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
system=[
{
"type": "text",
"text": "你是一个资深代码审查员,专注于 Python 和 Go……(很长的 prompt)",
"cache_control": {"type": "ephemeral"}, # 加这一行
}
],
messages=[{"role": "user", "content": code_snippet}],
)
cache hit 的 token 费用是原价的 10%,system prompt 长的话省一大截。
改完之后
| 项目 | 改前 | 改后 |
|------|------|------|
| 模型分布 | 100% Sonnet | 70% Haiku + 30% Sonnet |
| 超时重试浪费 | 偶发 ~15% | 几乎 0 |
| system prompt cache | 无 | hit rate ~80% |
| 月账单 | $217 | $58 |
不是什么黑魔法,就是把该省的地方系统性收了一遍。
一个可直接跑的 curl 验证
换完 base_url 先跑这个确认通了:
curl https://api.api2everything.xyz/v1/messages \
-H "x-api-key: 你的key" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-3-5",
"max_tokens": 64,
"messages": [{"role": "user", "content": "ping"}]
}'
能返回正常 JSON 就说明链路通了,后面直接切生产。
我用的中转是 api2everything.xyz,注册送 ¥1,够把上面这套流程验证一遍,余额也不会过期。支持 Claude / GPT / Gemini / DeepSeek 300+ 模型,格式全兼容 OpenAI,不用改 SDK。