Responses
The POST /v1/responses endpoint — built-in tools, reasoning parameters, and response structure.
Responses is OpenAI's next-generation unified API, and FIM Gate forwards it in the same format. Compared with Chat Completions, it replaces messages with a more flexible input, and adds built-in tools such as web search plus reasoning effort parameters.
Endpoint
| Item | Value |
|---|---|
| Method and path | POST /v1/responses |
| Base URL | https://api.gate.fim.ai/v1 |
| Authentication | Authorization: Bearer $FIM_API_KEY |
Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID, e.g. gpt-4.1, o3-mini |
input | string / array | Yes | Input content: a plain string, or an array of message objects with a role |
instructions | string | No | System-level instructions, similar to a system prompt |
stream | boolean | No | Set to true to enable SSE streaming events |
max_output_tokens | integer | No | Output token limit (including reasoning tokens) |
temperature | number | No | Sampling temperature, 0–2 |
top_p | number | No | Nucleus sampling threshold, 0–1 |
tools | array | No | Tool list, supporting built-in tools (e.g. {"type": "web_search_preview"}) and custom functions |
tool_choice | string / object | No | Tool calling strategy |
reasoning | object | No | Reasoning models only, e.g. {"effort": "low" | "medium" | "high"} to control thinking effort |
When input is a message array, the multimodal content block types are input_text and input_image (note these differ from Chat Completions' text / image_url naming).
Text request
curl https://api.gate.fim.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $FIM_API_KEY" \
-d '{
"model": "gpt-4.1",
"input": "Tell a bedtime story about a lighthouse in three sentences"
}'Response structure:
{
"id": "resp_abc123",
"object": "response",
"created_at": 1765500000,
"status": "completed",
"model": "gpt-4.1",
"output": [
{
"id": "msg_def456",
"type": "message",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "By the sea stood an old lighthouse...",
"annotations": []
}
]
}
],
"usage": {
"input_tokens": 18,
"output_tokens": 74,
"total_tokens": 92
}
}The generated text lives in content[].text of the type: "message" item within the output array.
status field
| Value | Meaning |
|---|---|
completed | Finished normally |
incomplete | Output was truncated; see incomplete_details.reason (e.g. max_output_tokens means the output limit was reached) |
failed | The request failed |
Image input
curl https://api.gate.fim.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $FIM_API_KEY" \
-d '{
"model": "gpt-4.1",
"input": [
{
"role": "user",
"content": [
{"type": "input_text", "text": "What is in this image?"},
{"type": "input_image", "image_url": "https://example.com/photo.jpg"}
]
}
]
}'Note that the image_url of an input_image block is a plain string, not a nested {"url": ...} object as in Chat Completions.
Built-in tool: web search
Declare the web_search_preview tool and the model can search the web on its own before answering:
curl https://api.gate.fim.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $FIM_API_KEY" \
-d '{
"model": "gpt-4.1",
"tools": [{"type": "web_search_preview"}],
"input": "Summarize one tech news story from today"
}'When search is used, the output array first contains a type: "web_search_call" item, followed by the message item with citation annotations.
OpenAI's deep-research models (such as o3-deep-research and o4-mini-deep-research) must be given a web search or MCP tool to run — otherwise the request fails immediately.
Reasoning parameters
For reasoning models, control thinking effort via reasoning.effort:
curl https://api.gate.fim.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $FIM_API_KEY" \
-d '{
"model": "o3-mini",
"input": "Which is larger, 9.11 or 9.9? Show your reasoning",
"reasoning": {"effort": "high"}
}'Higher effort means more thorough thinking and more reasoning tokens consumed. OpenAI reasoning models do not return the raw thinking text; reasoning token usage only shows up in usage.
Usage tips
- Be careful with
max_output_tokenson reasoning models: reasoning tokens count toward the limit, and a value that is too small yields an empty result withstatus: "incomplete"andincomplete_details.reasonset tomax_output_tokens. - Deep-research tasks take a long time — enable
streamand raise your client timeout accordingly.