FIM Gate 文档
使用指南

推理模型

通过 OpenAI 兼容接口控制 Claude 与 Gemini 思考过程的参数约定:thinking 开关、强度与预算。

本页的所有参数约定只在 OpenAI 兼容接口(https://api.gate.fim.ai/v1)中生效。如果你走 Claude 或 Gemini 的原生入口,请直接使用各家官方的思考参数。

Claude

max_tokens 默认值

通过 OpenAI 兼容接口调用 Claude 时,如果请求没有携带 max_tokens,FIM Gate 会自动填入该模型允许的最大值(例如 claude-3-7 系列为 128000),避免因缺少必填参数而报错。

用 #thinking 后缀开启思考

在模型名后追加 #thinking 即可开启 thinking 模式,此时 budget_tokens 自动取 max_tokens 的 80%:

curl https://api.gate.fim.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "claude-3-7-sonnet-20250219#thinking",
    "messages": [{"role": "user", "content": "24 点:用 3、3、8、8 凑出 24"}]
  }'

用 reasoning 参数精细控制

不改模型名也可以,在请求体中传入 reasoning 对象:

{
  "model": "claude-3-7-sonnet-20250219",
  "messages": [{"role": "user", "content": "..."}],
  "reasoning": {
    "effort": "high",
    "max_tokens": 2000
  }
}

两个字段都是可选的:

  • 传空对象 "reasoning": {} 即开启思考,效果与 #thinking 后缀相同。

  • effort 按比例分配思考预算,取值与 budget_tokens 的换算关系:

    effortbudget_tokens 占 max_tokens 的比例
    high80%
    medium50%
    low20%
  • reasoning.max_tokens 直接指定 budget_tokens 的具体数值,不能低于 1024,且必须小于请求的 max_tokens。

  • 两者同时传入时,reasoning.max_tokens 优先于 effort。

Gemini

Gemini 的约定与 Claude 相反:Gemini 2.5 系列默认开启思考,reasoning 参数主要用来限制或关闭它。

{
  "model": "gemini-2.5-flash",
  "messages": [{"role": "user", "content": "..."}],
  "reasoning": {
    "max_tokens": 2000
  }
}
  • reasoning.max_tokens 设为具体数值时,作为思考预算(budget_tokens)上限。
  • reasoning.max_tokens 设为 0,或只传空对象 "reasoning": {},表示关闭思考。
  • 如果客户端只能填模型名、无法自定义请求体,可直接使用别名 gemini-2.5-flash-nothinking,效果等同于关闭思考的 gemini-2.5-flash。

注意这一点与 Claude 的语义差异:空的 reasoning 对象在 Claude 上是开启思考,在 Gemini 上是关闭思考。