FIM Gate Docs
API Reference

Responses

The POST /v1/responses endpoint — built-in tools, reasoning parameters, and response structure.

Responses is OpenAI's next-generation unified API, and FIM Gate forwards it in the same format. Compared with Chat Completions, it replaces messages with a more flexible input, and adds built-in tools such as web search plus reasoning effort parameters.

Endpoint

ItemValue
Method and pathPOST /v1/responses
Base URLhttps://api.gate.fim.ai/v1
AuthenticationAuthorization: Bearer $FIM_API_KEY

Request parameters

ParameterTypeRequiredDescription
modelstringYesModel ID, e.g. gpt-4.1, o3-mini
inputstring / arrayYesInput content: a plain string, or an array of message objects with a role
instructionsstringNoSystem-level instructions, similar to a system prompt
streambooleanNoSet to true to enable SSE streaming events
max_output_tokensintegerNoOutput token limit (including reasoning tokens)
temperaturenumberNoSampling temperature, 0–2
top_pnumberNoNucleus sampling threshold, 0–1
toolsarrayNoTool list, supporting built-in tools (e.g. {"type": "web_search_preview"}) and custom functions
tool_choicestring / objectNoTool calling strategy
reasoningobjectNoReasoning models only, e.g. {"effort": "low" | "medium" | "high"} to control thinking effort

When input is a message array, the multimodal content block types are input_text and input_image (note these differ from Chat Completions' text / image_url naming).

Text request

curl https://api.gate.fim.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "gpt-4.1",
    "input": "Tell a bedtime story about a lighthouse in three sentences"
  }'

Response structure:

{
  "id": "resp_abc123",
  "object": "response",
  "created_at": 1765500000,
  "status": "completed",
  "model": "gpt-4.1",
  "output": [
    {
      "id": "msg_def456",
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "By the sea stood an old lighthouse...",
          "annotations": []
        }
      ]
    }
  ],
  "usage": {
    "input_tokens": 18,
    "output_tokens": 74,
    "total_tokens": 92
  }
}

The generated text lives in content[].text of the type: "message" item within the output array.

status field

ValueMeaning
completedFinished normally
incompleteOutput was truncated; see incomplete_details.reason (e.g. max_output_tokens means the output limit was reached)
failedThe request failed

Image input

curl https://api.gate.fim.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "gpt-4.1",
    "input": [
      {
        "role": "user",
        "content": [
          {"type": "input_text", "text": "What is in this image?"},
          {"type": "input_image", "image_url": "https://example.com/photo.jpg"}
        ]
      }
    ]
  }'

Note that the image_url of an input_image block is a plain string, not a nested {"url": ...} object as in Chat Completions.

Declare the web_search_preview tool and the model can search the web on its own before answering:

curl https://api.gate.fim.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "gpt-4.1",
    "tools": [{"type": "web_search_preview"}],
    "input": "Summarize one tech news story from today"
  }'

When search is used, the output array first contains a type: "web_search_call" item, followed by the message item with citation annotations.

OpenAI's deep-research models (such as o3-deep-research and o4-mini-deep-research) must be given a web search or MCP tool to run — otherwise the request fails immediately.

Reasoning parameters

For reasoning models, control thinking effort via reasoning.effort:

curl https://api.gate.fim.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "o3-mini",
    "input": "Which is larger, 9.11 or 9.9? Show your reasoning",
    "reasoning": {"effort": "high"}
  }'

Higher effort means more thorough thinking and more reasoning tokens consumed. OpenAI reasoning models do not return the raw thinking text; reasoning token usage only shows up in usage.

Usage tips

  • Be careful with max_output_tokens on reasoning models: reasoning tokens count toward the limit, and a value that is too small yields an empty result with status: "incomplete" and incomplete_details.reason set to max_output_tokens.
  • Deep-research tasks take a long time — enable stream and raise your client timeout accordingly.