Error Codes & FAQ
A quick reference for FIM Gate error codes, plus troubleshooting for common issues when calling LLM APIs.
This page summarizes the HTTP error codes returned by FIM Gate, along with the most common issues you may hit in practice and how to troubleshoot them.
Error code reference
| HTTP status | Error type | Meaning | Recommended action |
|---|---|---|---|
| 400 | Bad Request | Invalid request format or parameters | Check request body fields, types, and value ranges against the API docs |
| 401 | Unauthorized | API key missing or invalid | Confirm the request carries a valid FIM Gate API key and that the key has not been deleted or disabled |
| 404 | Not Found | Wrong request path | Check that the base URL and endpoint path are joined correctly (e.g. /v1/chat/completions) |
| 413 | Request Entity Too Large | Request body exceeds the size limit | Trim message content, or compress / split the input data |
| 429 | Too Many Requests | Rate limit triggered | Reduce request frequency and retry with exponential backoff; contact us for a higher quota if it keeps happening |
| 500 | Internal Server Error | Internal service error | Retry later; if it persists, contact support with the request details |
| 503 | Service Unavailable | Service temporarily unavailable (upstream maintenance or overload) | Retry later; clients should implement automatic retry and fallback logic |
Insufficient quota error, but your account balance is fine?
Each API key can have its own quota limit. Go to the FIM Gate Console and check whether the key you are using still has quota remaining; raise the key's quota or switch keys if needed.
"No available channel" / model unavailable?
- Check that the model name in your request is spelled correctly.
- Confirm the model is in the list of models currently available to your account — see the live model list at gate.fim.ai/pricing.
When a model name is misspelled or not enabled, the gateway returns an error response like:
{
"error": {
"code": "model_not_found",
"message": "No available channel (distributor) for model gpt-99-turbo in group default",
"type": "new_api_error"
}
}Seeing model_not_found means the problem is the model name itself, not any other field in the request body.
Reasoning models returning empty content?
If you get an empty response when calling a reasoning model such as GPT-5, first check whether you set max_tokens. Tokens consumed during reasoning also count toward max_tokens: if the limit is too small and reasoning tokens exhaust it first, the model stops before producing any visible output. Since OpenAI's reasoning models do not return their reasoning content, this shows up as an empty response.
Fix: remove max_tokens, or raise it.
With the Chat Completions API, you can check the finish_reason field to confirm whether the response was truncated by max_tokens:
| finish_reason | Meaning |
|---|---|
stop | Completed normally |
length | Hit the max_tokens limit |
content_filter | Content was filtered |
With the Responses API, check the status field instead: completed means it finished normally, incomplete means it stopped abnormally — in that case read incomplete_details.reason:
| reason | Meaning |
|---|---|
max_output_tokens | Hit the output token limit |
deep-research model requests failing?
Deep research models such as o3-deep-research and o4-mini-deep-research must be paired with a web search tool or an MCP tool to run; requests without tools fail immediately.
Using the web_search_preview tool:
curl -X POST 'https://api.gate.fim.ai/v1/responses' \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer $FIM_API_KEY" \
-d '{
"model": "o3-deep-research",
"stream": true,
"input": [
{ "role": "user", "content": "What were the major developments in AI over the past week?" }
],
"tools": [
{ "type": "web_search_preview" }
]
}'Using an MCP tool:
{
"model": "o3-deep-research",
"stream": true,
"input": [
{ "role": "user", "content": "Fetch and summarize this page: https://example.com" }
],
"tools": [
{
"type": "mcp",
"server_label": "my-mcp-server",
"server_url": "https://example.com/mcp/sse"
}
]
}GPT-5 series responding slowly?
The GPT-5 series are reasoning models, and OpenAI does not stream the reasoning process — output only starts after reasoning completes, so time-to-first-token is noticeably higher than with regular models.
You can trade reasoning depth for speed via reasoning.effort:
{
"model": "gpt-5",
"stream": true,
"input": "hi~",
"reasoning": {
"effort": "minimal"
}
}| effort | Description |
|---|---|
minimal | Least reasoning, fastest response; GPT-5 series only, not available on o-series models |
low | Less reasoning, faster response |
medium | Moderate reasoning, slower response |
high | Deepest reasoning, slowest response |
Gemini 2.5 Flash responding slowly?
Gemini 2.5 Flash is a hybrid reasoning model with thinking enabled by default. To get faster responses, you can turn thinking off:
- Via the OpenAI-compatible API: request
gemini-2.5-flash-nothinkingdirectly — FIM Gate provides this alias. - Via the native Gemini API: set
thinkingConfigingenerationConfig:
curl -X POST 'https://api.gate.fim.ai/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse' \
-H 'Content-Type: application/json' \
-H "x-goog-api-key: $FIM_API_KEY" \
-d '{
"contents": [
{ "role": "user", "parts": [{ "text": "hi~" }] }
],
"generationConfig": {
"thinkingConfig": {
"includeThoughts": false,
"thinkingBudget": 0
}
}
}'Reasoning models frequently timing out / failing?
Reasoning models can take a very long time on complex tasks (over 2,000 seconds in extreme cases), so default client timeouts easily cause disconnects. Recommendations:
- Prefer
stream(streaming) mode — continuously receiving data avoids idle timeouts. - If you must use non-streaming requests, raise the client's request timeout.
- If traffic goes through a proxy or relay, make sure the timeout and idle-disconnect settings of every hop on the path are large enough too.
Why can't the model say which model it is?
This is not a FIM Gate forwarding issue — it is a general phenomenon with large language models:
- Training data predates release: the training corpus does not contain the model's own name, so the model has no way to "know" what it is called.
- No self-awareness: a model is fundamentally a next-token predictor; answers about its identity must be injected explicitly.
Official web apps like ChatGPT and Claude answer correctly because their system prompts state the identity. If your application also needs the model to report its identity correctly, add something like this to the system prompt:
You are GPT-4, a large language model trained by OpenAI.