省钱实战 · 阅读约 4 分钟

月花 200 刀调 Claude 的我,改了一行代码后现在只花 60

月花 200 刀调 Claude 的我,改了一行代码后现在只花 60

结论先说:把 base_url 换掉,账单直接降到三分之一。


痛点

我有个跑了半年的 code review bot,后端全是 Claude Sonnet。上个月账单 $217,看着有点难绷。

细拆了一下,问题很集中:

1. 所有请求都打 Sonnet,不管是"格式化这段代码"还是"解释这个架构",统一顶配

2. 没做任何缓存,同一段 system prompt 每次都重新计费

3. 直连 api.anthropic.com,国内偶尔超时重试,token 白费

其实第 3 条是最亏的——超时之后客户端自动重试,等于同一个请求付了两次钱,但我只拿到一次结果。


怎么改

第一步:换 base_url,解决超时重试问题

原来的代码:


import anthropic



client = anthropic.Anthropic(api_key="sk-ant-xxx")

改成:


import anthropic



client = anthropic.Anthropic(

    api_key="你的中转 key",

    base_url="https://api.api2everything.xyz/v1",

)

接口格式完全兼容,其他代码一行不动。用 OpenAI SDK 也一样,改 base_url 就完事。

第二步:按任务分流模型

不是所有任务都需要 Sonnet。我按复杂度拆成两档:


def get_model(task_type: str) -> str:

    # 简单任务:格式化、注释、重命名

    if task_type in ("format", "comment", "rename"):

        return "claude-haiku-3-5"   # 比 Sonnet 便宜 ~20x

    # 复杂任务:架构分析、bug 定位

    return "claude-sonnet-4-5"



def review(code: str, task_type: str = "general"):

    resp = client.messages.create(

        model=get_model(task_type),

        max_tokens=1024,

        messages=[{"role": "user", "content": code}],

    )

    return resp.content[0].text

第三步:system prompt 走 cache(Anthropic 原生支持)

长 system prompt 每次重新传很费钱,加一个 cache_control 就能复用:


response = client.messages.create(

    model="claude-sonnet-4-5",

    max_tokens=1024,

    system=[

        {

            "type": "text",

            "text": "你是一个资深代码审查员,专注于 Python 和 Go……(很长的 prompt)",

            "cache_control": {"type": "ephemeral"},  # 加这一行

        }

    ],

    messages=[{"role": "user", "content": code_snippet}],

)

cache hit 的 token 费用是原价的 10%,system prompt 长的话省一大截。


改完之后

| 项目 | 改前 | 改后 |

|------|------|------|

| 模型分布 | 100% Sonnet | 70% Haiku + 30% Sonnet |

| 超时重试浪费 | 偶发 ~15% | 几乎 0 |

| system prompt cache | 无 | hit rate ~80% |

| 月账单 | $217 | $58 |

不是什么黑魔法,就是把该省的地方系统性收了一遍。


一个可直接跑的 curl 验证

换完 base_url 先跑这个确认通了:


curl https://api.api2everything.xyz/v1/messages \

  -H "x-api-key: 你的key" \

  -H "anthropic-version: 2023-06-01" \

  -H "content-type: application/json" \

  -d '{

    "model": "claude-haiku-3-5",

    "max_tokens": 64,

    "messages": [{"role": "user", "content": "ping"}]

  }'

能返回正常 JSON 就说明链路通了,后面直接切生产。


我用的中转是 api2everything.xyz,注册送 ¥1,够把上面这套流程验证一遍,余额也不会过期。支持 Claude / GPT / Gemini / DeepSeek 300+ 模型,格式全兼容 OpenAI,不用改 SDK。