Gemini Native API
Call Gemini models using the native Google Gemini API format.
When you need field-for-field compatibility with the Google Gemini API (GenAI style), use FIM Gate's Gemini native entry point. Gemini's interface differs significantly from OpenAI's: the model name goes in the URL path, and the request body uses a contents / parts structure. If you don't need that compatibility, consider the OpenAI-compatible API instead.
Endpoint
| Item | Value |
|---|---|
| Base URL | https://api.gate.fim.ai |
| Non-streaming | POST /v1beta/models/{model}:generateContent |
| Streaming | POST /v1beta/models/{model}:streamGenerateContent?alt=sse |
| Authentication | x-goog-api-key: $FIM_API_KEY |
Replace {model} with a model ID such as gemini-2.5-flash or gemini-2.5-pro. Streaming requests must include the query parameter alt=sse — without it, you get chunked JSON array fragments instead of standard SSE.
Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
contents | array | Yes | Array of conversation contents, each with a role (user / model) and parts |
contents[].parts | array | Yes | Array of content blocks; text uses {"text": "..."}, inline images use inline_data (with mime_type and base64 data) |
systemInstruction | object | No | System instruction, structured as {"parts": [{"text": "..."}]} |
generationConfig | object | No | Generation parameter set, see table below |
safetySettings | array | No | Safety filter threshold configuration |
tools | array | No | Tool declarations; function calling uses functionDeclarations |
Common generationConfig fields:
| Field | Type | Description |
|---|---|---|
temperature | number | Sampling temperature |
topP | number | Nucleus sampling threshold |
topK | integer | Candidate cutoff count |
maxOutputTokens | integer | Output token limit |
stopSequences | array | Stop sequences |
Note that the Gemini native format uses camelCase field names (maxOutputTokens), unlike OpenAI's snake_case naming.
Non-streaming request
curl "https://api.gate.fim.ai/v1beta/models/gemini-2.5-flash:generateContent" \
-H "Content-Type: application/json" \
-H "x-goog-api-key: $FIM_API_KEY" \
-d '{
"contents": [
{
"role": "user",
"parts": [{"text": "Why is seawater salty?"}]
}
],
"generationConfig": {
"temperature": 0.7,
"maxOutputTokens": 1024
}
}'Response structure:
{
"candidates": [
{
"content": {
"role": "model",
"parts": [
{"text": "The salt in seawater comes mainly from rock weathering..."}
]
},
"finishReason": "STOP",
"index": 0
}
],
"usageMetadata": {
"promptTokenCount": 9,
"candidatesTokenCount": 182,
"totalTokenCount": 191
},
"modelVersion": "gemini-2.5-flash"
}The generated text is in candidates[0].content.parts[].text.
finishReason values
| Value | Meaning |
|---|---|
STOP | Finished normally |
MAX_TOKENS | Reached the maxOutputTokens limit |
SAFETY | Triggered safety filtering |
RECITATION | Stopped due to recitation detection |
Streaming request (SSE)
curl "https://api.gate.fim.ai/v1beta/models/gemini-2.5-flash:streamGenerateContent?alt=sse" \
-H "Content-Type: application/json" \
-H "x-goog-api-key: $FIM_API_KEY" \
-d '{
"contents": [
{
"role": "user",
"parts": [{"text": "Write a 50-word product introduction"}]
}
]
}'The data of each SSE event is a JSON fragment with the same shape as the non-streaming response. The incremental text is likewise in candidates[0].content.parts[].text — just concatenate it event by event.
Using the Google GenAI SDK
pip install -U google-genaifrom google import genai
from google.genai import types
client = genai.Client(
api_key="$FIM_API_KEY",
http_options=types.HttpOptions(base_url="https://api.gate.fim.ai"),
)
response = client.models.generate_content(
model="gemini-2.5-flash",
contents="Hello",
)
print(response.text)Point the SDK's requests at FIM Gate via http_options.base_url; everything else works the same as the official SDK.