Reasoning Models
Parameter conventions for controlling Claude and Gemini thinking through the OpenAI-compatible API, including the thinking toggle, effort, and budget.
All parameter conventions on this page apply only to the OpenAI-compatible API (https://api.gate.fim.ai/v1). If you use the Claude or Gemini native entry points, use each provider's official thinking parameters instead.
Claude
max_tokens default
When calling Claude through the OpenAI-compatible API, if the request omits max_tokens, FIM Gate automatically fills in the model's maximum allowed value (e.g. 128000 for the claude-3-7 series), so the request doesn't fail on a missing required parameter.
Enable thinking with the #thinking suffix
Append #thinking to the model name to enable thinking mode; budget_tokens is then automatically set to 80% of max_tokens:
curl https://api.gate.fim.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $FIM_API_KEY" \
-d '{
"model": "claude-3-7-sonnet-20250219#thinking",
"messages": [{"role": "user", "content": "24 game: make 24 from 3, 3, 8, 8"}]
}'Fine-grained control with the reasoning parameter
You can also keep the model name unchanged and pass a reasoning object in the request body:
{
"model": "claude-3-7-sonnet-20250219",
"messages": [{"role": "user", "content": "..."}],
"reasoning": {
"effort": "high",
"max_tokens": 2000
}
}Both fields are optional:
-
Passing an empty object
"reasoning": {}enables thinking, equivalent to the#thinkingsuffix. -
effortallocates the thinking budget proportionally. The mapping tobudget_tokens:effort budget_tokens as a share of max_tokens high80% medium50% low20% -
reasoning.max_tokenssets the exactbudget_tokensvalue. It must be at least 1024 and less than the request'smax_tokens. -
When both are provided,
reasoning.max_tokenstakes precedence overeffort.
Gemini
Gemini's convention is the opposite of Claude's: the Gemini 2.5 series has thinking enabled by default, and the reasoning parameter is mainly used to limit or disable it.
{
"model": "gemini-2.5-flash",
"messages": [{"role": "user", "content": "..."}],
"reasoning": {
"max_tokens": 2000
}
}- Setting
reasoning.max_tokensto a specific value caps the thinking budget (budget_tokens). - Setting
reasoning.max_tokensto0, or passing just an empty object"reasoning": {}, disables thinking. - If your client only accepts a model name and cannot customize the request body, use the alias
gemini-2.5-flash-nothinking, which behaves likegemini-2.5-flashwith thinking disabled.
Note this semantic difference from Claude: an empty reasoning object enables thinking on Claude but disables it on Gemini.