API Reference
Audio API (TTS / STT)
Text-to-speech and speech-to-text through the OpenAI-compatible API.
FIM Gate provides OpenAI-compatible audio endpoints covering two scenarios:
- Text-to-speech (TTS): send text, receive an audio file
- Speech-to-text (STT): upload an audio file, receive the transcribed text
Basics
| Item | Value |
|---|---|
| Base URL | https://api.gate.fim.ai/v1 |
| Authentication | Authorization: Bearer $FIM_API_KEY header |
Text-to-speech (Speech)
POST /v1/audio/speech
The request body is JSON, and the endpoint returns a binary audio stream directly — redirect it to a file to save it.
Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | TTS model ID, e.g. tts-1 |
input | string | Yes | The text to synthesize; Chinese is supported |
voice | string | Yes | Voice name, e.g. alloy |
response_format | string | No | Output format, e.g. mp3, wav |
speed | number | No | Speech speed multiplier |
Currently available TTS models: tts-1, tts-1-hd, gpt-4o-mini-tts, gemini-2.5-flash-preview-tts, gemini-2.5-pro-preview-tts. See gate.fim.ai/pricing for the full list.
Example request
curl https://api.gate.fim.ai/v1/audio/speech \
-H "Authorization: Bearer $FIM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "Welcome to FIM Gate — one line of config to access models from many providers.",
"voice": "alloy"
}' \
--output welcome.mp3On success, welcome.mp3 is created in the current directory.
Speech-to-text (Transcriptions)
POST /v1/audio/transcriptions
The request body is multipart/form-data with an uploaded audio file.
Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
file | file | Yes | The audio file to transcribe (e.g. mp3, wav, m4a) |
model | string | Yes | Transcription model ID, e.g. whisper-1 |
language | string | No | Audio language (ISO-639-1, e.g. zh); improves accuracy |
prompt | string | No | Prompt text to guide transcription style or proper nouns |
response_format | string | No | Response format, e.g. json, text, srt |
temperature | number | No | Sampling temperature, 0 to 1 |
Currently available transcription models: whisper-1. See gate.fim.ai/pricing for the full list.
Example request
curl https://api.gate.fim.ai/v1/audio/transcriptions \
-H "Authorization: Bearer $FIM_API_KEY" \
-H "Content-Type: multipart/form-data" \
-F file=@./meeting.mp3 \
-F model=whisper-1Example response
{
"text": "This is the transcribed text content."
}FAQ
- TTS returns a binary stream — save it to disk with
--outputor an equivalent; don't try to parse it as JSON. - When uploading large files to STT, watch your client timeout settings and split the audio first if necessary.