FIM Gate Docs
Guide

Reasoning Models

Parameter conventions for controlling Claude and Gemini thinking through the OpenAI-compatible API, including the thinking toggle, effort, and budget.

All parameter conventions on this page apply only to the OpenAI-compatible API (https://api.gate.fim.ai/v1). If you use the Claude or Gemini native entry points, use each provider's official thinking parameters instead.

Claude

max_tokens default

When calling Claude through the OpenAI-compatible API, if the request omits max_tokens, FIM Gate automatically fills in the model's maximum allowed value (e.g. 128000 for the claude-3-7 series), so the request doesn't fail on a missing required parameter.

Enable thinking with the #thinking suffix

Append #thinking to the model name to enable thinking mode; budget_tokens is then automatically set to 80% of max_tokens:

curl https://api.gate.fim.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "claude-3-7-sonnet-20250219#thinking",
    "messages": [{"role": "user", "content": "24 game: make 24 from 3, 3, 8, 8"}]
  }'

Fine-grained control with the reasoning parameter

You can also keep the model name unchanged and pass a reasoning object in the request body:

{
  "model": "claude-3-7-sonnet-20250219",
  "messages": [{"role": "user", "content": "..."}],
  "reasoning": {
    "effort": "high",
    "max_tokens": 2000
  }
}

Both fields are optional:

  • Passing an empty object "reasoning": {} enables thinking, equivalent to the #thinking suffix.

  • effort allocates the thinking budget proportionally. The mapping to budget_tokens:

    effortbudget_tokens as a share of max_tokens
    high80%
    medium50%
    low20%
  • reasoning.max_tokens sets the exact budget_tokens value. It must be at least 1024 and less than the request's max_tokens.

  • When both are provided, reasoning.max_tokens takes precedence over effort.

Gemini

Gemini's convention is the opposite of Claude's: the Gemini 2.5 series has thinking enabled by default, and the reasoning parameter is mainly used to limit or disable it.

{
  "model": "gemini-2.5-flash",
  "messages": [{"role": "user", "content": "..."}],
  "reasoning": {
    "max_tokens": 2000
  }
}
  • Setting reasoning.max_tokens to a specific value caps the thinking budget (budget_tokens).
  • Setting reasoning.max_tokens to 0, or passing just an empty object "reasoning": {}, disables thinking.
  • If your client only accepts a model name and cannot customize the request body, use the alias gemini-2.5-flash-nothinking, which behaves like gemini-2.5-flash with thinking disabled.

Note this semantic difference from Claude: an empty reasoning object enables thinking on Claude but disables it on Gemini.