FIM Gate Docs
API Reference

Gemini Native API

Call Gemini models using the native Google Gemini API format.

When you need field-for-field compatibility with the Google Gemini API (GenAI style), use FIM Gate's Gemini native entry point. Gemini's interface differs significantly from OpenAI's: the model name goes in the URL path, and the request body uses a contents / parts structure. If you don't need that compatibility, consider the OpenAI-compatible API instead.

Endpoint

ItemValue
Base URLhttps://api.gate.fim.ai
Non-streamingPOST /v1beta/models/{model}:generateContent
StreamingPOST /v1beta/models/{model}:streamGenerateContent?alt=sse
Authenticationx-goog-api-key: $FIM_API_KEY

Replace {model} with a model ID such as gemini-2.5-flash or gemini-2.5-pro. Streaming requests must include the query parameter alt=sse — without it, you get chunked JSON array fragments instead of standard SSE.

Request parameters

ParameterTypeRequiredDescription
contentsarrayYesArray of conversation contents, each with a role (user / model) and parts
contents[].partsarrayYesArray of content blocks; text uses {"text": "..."}, inline images use inline_data (with mime_type and base64 data)
systemInstructionobjectNoSystem instruction, structured as {"parts": [{"text": "..."}]}
generationConfigobjectNoGeneration parameter set, see table below
safetySettingsarrayNoSafety filter threshold configuration
toolsarrayNoTool declarations; function calling uses functionDeclarations

Common generationConfig fields:

FieldTypeDescription
temperaturenumberSampling temperature
topPnumberNucleus sampling threshold
topKintegerCandidate cutoff count
maxOutputTokensintegerOutput token limit
stopSequencesarrayStop sequences

Note that the Gemini native format uses camelCase field names (maxOutputTokens), unlike OpenAI's snake_case naming.

Non-streaming request

curl "https://api.gate.fim.ai/v1beta/models/gemini-2.5-flash:generateContent" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: $FIM_API_KEY" \
  -d '{
    "contents": [
      {
        "role": "user",
        "parts": [{"text": "Why is seawater salty?"}]
      }
    ],
    "generationConfig": {
      "temperature": 0.7,
      "maxOutputTokens": 1024
    }
  }'

Response structure:

{
  "candidates": [
    {
      "content": {
        "role": "model",
        "parts": [
          {"text": "The salt in seawater comes mainly from rock weathering..."}
        ]
      },
      "finishReason": "STOP",
      "index": 0
    }
  ],
  "usageMetadata": {
    "promptTokenCount": 9,
    "candidatesTokenCount": 182,
    "totalTokenCount": 191
  },
  "modelVersion": "gemini-2.5-flash"
}

The generated text is in candidates[0].content.parts[].text.

finishReason values

ValueMeaning
STOPFinished normally
MAX_TOKENSReached the maxOutputTokens limit
SAFETYTriggered safety filtering
RECITATIONStopped due to recitation detection

Streaming request (SSE)

curl "https://api.gate.fim.ai/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse" \
  -H "Content-Type: application/json" \
  -H "x-goog-api-key: $FIM_API_KEY" \
  -d '{
    "contents": [
      {
        "role": "user",
        "parts": [{"text": "Write a 50-word product introduction"}]
      }
    ]
  }'

The data of each SSE event is a JSON fragment with the same shape as the non-streaming response. The incremental text is likewise in candidates[0].content.parts[].text — just concatenate it event by event.

Using the Google GenAI SDK

pip install -U google-genai
from google import genai
from google.genai import types

client = genai.Client(
    api_key="$FIM_API_KEY",
    http_options=types.HttpOptions(base_url="https://api.gate.fim.ai"),
)

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Hello",
)
print(response.text)

Point the SDK's requests at FIM Gate via http_options.base_url; everything else works the same as the official SDK.