FIM Gate Docs
API Reference

Chat Completions

Parameters, streaming, and tool calling for POST /v1/chat/completions.

Chat Completions is the most widely used chat endpoint on FIM Gate. It follows the OpenAI Chat Completions specification and supports text conversations, image input, SSE streaming, and function calling.

Endpoint

ItemValue
Method and pathPOST /v1/chat/completions
Base URLhttps://api.gate.fim.ai/v1
AuthenticationAuthorization: Bearer $FIM_API_KEY

Request parameters

ParameterTypeRequiredDescription
modelstringYesModel ID, e.g. gpt-4o-mini, gpt-4.1, claude-sonnet-4-5
messagesarrayYesArray of conversation messages, each with a role (system / user / assistant / tool) and content
streambooleanNoSet to true to enable SSE streaming; defaults to false
max_tokensintegerNoOutput token limit. Note that reasoning models' thinking tokens count toward this limit — setting it too low can result in an empty response
temperaturenumberNoSampling temperature, 0–2; higher values produce more varied output
top_pnumberNoNucleus sampling threshold, 0–1; adjust either this or temperature, not both
nintegerNoNumber of candidate completions to generate per request, defaults to 1
stopstring / arrayNoStop sequences — generation halts when one is matched; up to 4
presence_penaltynumberNo-2.0–2.0; positive values encourage introducing new topics
frequency_penaltynumberNo-2.0–2.0; positive values discourage repeated wording
response_formatobjectNoSet {"type": "json_object"} to force the output to be valid JSON
toolsarrayNoList of functions the model may call, see below
tool_choicestring / objectNoTool calling strategy: auto, none, required, or a specific function
userstringNoEnd-user identifier, useful for auditing and abuse tracking

The content in messages can be either a plain string or an array of content blocks (for multimodal input, see below).

Text conversation

curl https://api.gate.fim.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [
      {"role": "system", "content": "You are a concise assistant"},
      {"role": "user", "content": "Explain what idempotency means"}
    ]
  }'

Response structure:

{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1765500000,
  "model": "gpt-4o-mini",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Idempotency means performing an operation once has the same effect as performing it multiple times..."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 28,
    "completion_tokens": 96,
    "total_tokens": 124
  }
}

finish_reason values

ValueMeaning
stopThe model finished naturally or matched a stop sequence
lengthOutput was truncated at the max_tokens limit
tool_callsThe model requested a tool call
content_filterContent was filtered by safety policy

Image input (multimodal)

Write content as an array of content blocks, mixing text and image_url blocks. Images can be public URLs or inline data in data:image/...;base64, form.

curl https://api.gate.fim.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "gpt-4.1",
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "text", "text": "Describe what is in this image"},
          {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
        ]
      }
    ],
    "max_tokens": 300
  }'

Streaming (SSE)

Add "stream": true to the request body and the response becomes an SSE event stream. Each event is a chat.completion.chunk object, with the incremental text in choices[0].delta.content.

curl https://api.gate.fim.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "gpt-4o-mini",
    "stream": true,
    "messages": [
      {"role": "user", "content": "Write a short poem about autumn"}
    ]
  }'

Example chunk structure:

{
  "id": "chatcmpl-abc123",
  "object": "chat.completion.chunk",
  "created": 1765500000,
  "model": "gpt-4o-mini",
  "choices": [
    {
      "index": 0,
      "delta": {"content": "Falling leaves"},
      "finish_reason": null
    }
  ]
}

Detecting the end: the last meaningful chunk carries finish_reason: "stop" (or another stop reason), followed by data: [DONE].

Function calling (Tools)

Declare functions and their JSON Schema parameters via tools. When the model decides a call is needed, it returns tool_calls in the response; your code executes the function and sends the result back as a role: "tool" message.

curl https://api.gate.fim.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "gpt-4.1",
    "messages": [
      {"role": "user", "content": "What is the weather like in Shanghai right now?"}
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "get_current_weather",
          "description": "Look up the current weather for a given city",
          "parameters": {
            "type": "object",
            "properties": {
              "city": {"type": "string", "description": "City name, e.g. Shanghai"},
              "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
            },
            "required": ["city"]
          }
        }
      }
    ],
    "tool_choice": "auto"
  }'

Response fragment when the model decides to call a tool:

{
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_xyz789",
            "type": "function",
            "function": {
              "name": "get_current_weather",
              "arguments": "{\"city\": \"Shanghai\", \"unit\": \"celsius\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ]
}

Usage tips

  • For reasoning models (such as the o-series), either omit max_tokens or set it generously: thinking tokens count toward the limit, and a value that is too small yields empty content with finish_reason set to length.
  • In long conversations, keep an eye on the total length of messages — exceeding the model's context window returns a 400.