FIM Gate Docs

Error Codes & FAQ

A quick reference for FIM Gate error codes, plus troubleshooting for common issues when calling LLM APIs.

This page summarizes the HTTP error codes returned by FIM Gate, along with the most common issues you may hit in practice and how to troubleshoot them.

Error code reference

HTTP statusError typeMeaningRecommended action
400Bad RequestInvalid request format or parametersCheck request body fields, types, and value ranges against the API docs
401UnauthorizedAPI key missing or invalidConfirm the request carries a valid FIM Gate API key and that the key has not been deleted or disabled
404Not FoundWrong request pathCheck that the base URL and endpoint path are joined correctly (e.g. /v1/chat/completions)
413Request Entity Too LargeRequest body exceeds the size limitTrim message content, or compress / split the input data
429Too Many RequestsRate limit triggeredReduce request frequency and retry with exponential backoff; contact us for a higher quota if it keeps happening
500Internal Server ErrorInternal service errorRetry later; if it persists, contact support with the request details
503Service UnavailableService temporarily unavailable (upstream maintenance or overload)Retry later; clients should implement automatic retry and fallback logic

Insufficient quota error, but your account balance is fine?

Each API key can have its own quota limit. Go to the FIM Gate Console and check whether the key you are using still has quota remaining; raise the key's quota or switch keys if needed.

"No available channel" / model unavailable?

  1. Check that the model name in your request is spelled correctly.
  2. Confirm the model is in the list of models currently available to your account — see the live model list at gate.fim.ai/pricing.

When a model name is misspelled or not enabled, the gateway returns an error response like:

{
  "error": {
    "code": "model_not_found",
    "message": "No available channel (distributor) for model gpt-99-turbo in group default",
    "type": "new_api_error"
  }
}

Seeing model_not_found means the problem is the model name itself, not any other field in the request body.

Reasoning models returning empty content?

If you get an empty response when calling a reasoning model such as GPT-5, first check whether you set max_tokens. Tokens consumed during reasoning also count toward max_tokens: if the limit is too small and reasoning tokens exhaust it first, the model stops before producing any visible output. Since OpenAI's reasoning models do not return their reasoning content, this shows up as an empty response.

Fix: remove max_tokens, or raise it.

With the Chat Completions API, you can check the finish_reason field to confirm whether the response was truncated by max_tokens:

finish_reasonMeaning
stopCompleted normally
lengthHit the max_tokens limit
content_filterContent was filtered

With the Responses API, check the status field instead: completed means it finished normally, incomplete means it stopped abnormally — in that case read incomplete_details.reason:

reasonMeaning
max_output_tokensHit the output token limit

deep-research model requests failing?

Deep research models such as o3-deep-research and o4-mini-deep-research must be paired with a web search tool or an MCP tool to run; requests without tools fail immediately.

Using the web_search_preview tool:

curl -X POST 'https://api.gate.fim.ai/v1/responses' \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $FIM_API_KEY" \
  -d '{
    "model": "o3-deep-research",
    "stream": true,
    "input": [
      { "role": "user", "content": "What were the major developments in AI over the past week?" }
    ],
    "tools": [
      { "type": "web_search_preview" }
    ]
  }'

Using an MCP tool:

{
  "model": "o3-deep-research",
  "stream": true,
  "input": [
    { "role": "user", "content": "Fetch and summarize this page: https://example.com" }
  ],
  "tools": [
    {
      "type": "mcp",
      "server_label": "my-mcp-server",
      "server_url": "https://example.com/mcp/sse"
    }
  ]
}

GPT-5 series responding slowly?

The GPT-5 series are reasoning models, and OpenAI does not stream the reasoning process — output only starts after reasoning completes, so time-to-first-token is noticeably higher than with regular models.

You can trade reasoning depth for speed via reasoning.effort:

{
  "model": "gpt-5",
  "stream": true,
  "input": "hi~",
  "reasoning": {
    "effort": "minimal"
  }
}
effortDescription
minimalLeast reasoning, fastest response; GPT-5 series only, not available on o-series models
lowLess reasoning, faster response
mediumModerate reasoning, slower response
highDeepest reasoning, slowest response

Gemini 2.5 Flash responding slowly?

Gemini 2.5 Flash is a hybrid reasoning model with thinking enabled by default. To get faster responses, you can turn thinking off:

  • Via the OpenAI-compatible API: request gemini-2.5-flash-nothinking directly — FIM Gate provides this alias.
  • Via the native Gemini API: set thinkingConfig in generationConfig:
curl -X POST 'https://api.gate.fim.ai/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse' \
  -H 'Content-Type: application/json' \
  -H "x-goog-api-key: $FIM_API_KEY" \
  -d '{
    "contents": [
      { "role": "user", "parts": [{ "text": "hi~" }] }
    ],
    "generationConfig": {
      "thinkingConfig": {
        "includeThoughts": false,
        "thinkingBudget": 0
      }
    }
  }'

Reasoning models frequently timing out / failing?

Reasoning models can take a very long time on complex tasks (over 2,000 seconds in extreme cases), so default client timeouts easily cause disconnects. Recommendations:

  1. Prefer stream (streaming) mode — continuously receiving data avoids idle timeouts.
  2. If you must use non-streaming requests, raise the client's request timeout.
  3. If traffic goes through a proxy or relay, make sure the timeout and idle-disconnect settings of every hop on the path are large enough too.

Why can't the model say which model it is?

This is not a FIM Gate forwarding issue — it is a general phenomenon with large language models:

  • Training data predates release: the training corpus does not contain the model's own name, so the model has no way to "know" what it is called.
  • No self-awareness: a model is fundamentally a next-token predictor; answers about its identity must be injected explicitly.

Official web apps like ChatGPT and Claude answer correctly because their system prompts state the identity. If your application also needs the model to report its identity correctly, add something like this to the system prompt:

You are GPT-4, a large language model trained by OpenAI.