Model Gateway
The Model Gateway is the single access point for every model on the platform. It exposes a supported subset of OpenAI-compatible model APIs together with platform extensions. Whether a model is deployed in your cluster or provided by an external service, clients use one base URL and the gateway routes the request to the right place, applying authorization and usage accounting along the way.
Point any OpenAI-compatible client at the base URL https://api.<your-domain>/v1 and pass an API key. The portal shows your platform's own URL as API URL at the bottom of the sidebar, with a button to copy it; add /v1 to it:
curl -sS "https://api.<your-domain>/v1/chat/completions" \
-H "Authorization: Bearer cm_api_…" \
-H "Content-Type: application/json" \
-d '{"model": "YOUR_MODEL_ID", "messages": [{"role": "user", "content": "Hello"}]}'
Endpoints
| Endpoint | Purpose |
|---|---|
POST /v1/chat/completions | OpenAI-compatible chat completions |
POST /v1/audio/transcriptions | Audio transcription — speech to text (guide) |
POST /v1/embeddings | OpenAI-compatible embeddings |
POST /v1/responses | OpenAI-compatible Responses API (guide) |
GET /v1/models | List available models |
OpenAI clients can call the documented models, Chat Completions, Embeddings, and Responses operations after you set the base URL and key. The gateway does not expose every operation or parameter in the OpenAI API; the operations and behaviors documented here define the supported public surface.
Authentication
Every request needs an API key as a bearer token, and the key must have invoke granted on the model — or on an alias that points at it. A request for a model it can't invoke returns 403.
For external models you still authenticate with your platform API key. The gateway holds the provider's credentials and attaches them when forwarding, so provider keys never live in your client.
Calling a model
Pass the model's id — exactly as returned by GET /v1/models — in the model field. The ID can identify a
model directly or a model alias: a stable public model ID that an administrator can point at another model.
If your API key is granted on an alias, keep sending the same ID after the alias is repointed. Requests go to the new model without requiring changes to the client or API key. See API Keys for choosing between an alias grant and a direct model grant. Administrators can manage aliases under Model Aliases.
The gateway resolves the public model ID, applies authorization and accounting, and may validate or transform the request before it reaches the backend. The following features are available where the selected model and backend support them:
- Streaming — set
"stream": truefor Server-Sent Events; add"stream_options": {"include_usage": true}for a final token-usage event. - Structured output — use
response_format(json_schemaorjson_object) to constrain output to a schema, where the backend supports it. - Image input — send
image_urlcontent parts to any model that listsimageamong its input modalities (see below).
Successful chat and embeddings responses carry the usual usage token counts.
For practical Responses examples, including streaming and stored conversations, see the Responses API guide.
Instructions on the Responses API
For a direct model request, POST /v1/responses accepts instructions as a string or null. Use it to give the model
a system prompt for that request.
Before each model turn, the Model Gateway adds context describing the tools the model may currently call.
With chat_mode: true, it also adds Chat formatting guidance and the current UTC date. Direct-model Chat requests
also receive the tenant's assistant instructions; preset agents do not. See Chat mode.
Requests that omit chat_mode or set it to false receive no automatic UTC date. Earlier platform versions added
it to every Responses request. If your integration relies on that date, send chat_mode: true to enable the full
Chat bundle, or include the date in your own direct-model instructions or preset-agent instructions.
This additional context is used only as model input. It is not saved in platform conversation history or returned in
client-facing response snapshots. The response's instructions field contains the original string you sent, or null
when you omitted the field or sent null.
Preset agents use the system prompt saved in their configuration. Do not send instructions when invoking a preset
agent; the field is not permitted on preset-agent requests. A preset-agent response always contains
"instructions": null, so its configured prompt is not exposed.
Any other JSON value for instructions returns 400 with the message
instructions must be a string or null.
chat_mode accepts only true or false. Any other JSON value, including null, returns 400 with the message
field chat_mode must be a boolean.
Preset-agent output guardrails
A preset agent can check its final answer with
up to four judge models before releasing it. The checks run in order and stop at the first block or error. They
apply whether the agent is called through Chat or POST /v1/responses. Direct model calls do not inherit an agent's
guardrails.
Each judge receives the configured policy, the current request's user text, and the final assistant answer text. This does not include the full conversation history, tool results, retrieval evidence, or multimodal content. Intermediate tool turns are not judged as final answers.
- Allowed: the gateway returns the selected final answer text or refusal.
- Blocked: the gateway returns a completed response containing the guardrail's block message instead.
- Judge error: the request fails without releasing the protected answer. If streaming has already started, the failure is delivered in the stream.
Every output guardrail keeps intermediate output private for the entire request. Users see no incremental answer text, reasoning or tool progress while generation and judging run. Streaming clients receive only lifecycle events and content-free named keepalives during this wait. After the checks, the gateway immediately serializes the selected final answer or block message into SSE events without pacing, so the result appears all at once. JSON uses the same selected result. Intermediate assistant text, reasoning, tool arguments/results, backend item IDs and extra model fields stay private. Final answer text/refusal, gateway-owned IDs/status, usage and safe errors remain. Annotations, citations and logprobs are omitted because the judge does not evaluate their exposed content.
The wire remains named Responses events followed by [DONE], without library heartbeat comments or retry frames.
Unguarded streaming continues to deliver answer content as it is generated.
Withholding keeps full execution context only within the request. When conversation storage is enabled, only user input and the selected final answer are persisted and available through history or forks. Later requests may repeat tool calls or lose details needed for follow-up questions. There is no private persistent history. Public content capture excludes private tool and judge content even when tenant content capture is enabled. Caller-delegated tools are rejected before execution on guarded requests. Gateway-executed tools still work, and trusted gateway OAuth recovery uses the existing error envelope without releasing a tool result.
A rejected final answer is not saved; its replacement is saved and used on subsequent turns. Guardrails do not undo stored input or actions already performed by tools. Adding guardrails does not erase older stored history. Base-model tokens and judge-model tokens are attributed to their respective targets in usage accounting.
Guarded SSE and JSON requests share the following limits per gateway process:
| Limit | Value |
|---|---|
| Concurrent admitted requests | 4 |
| Backend SSE wire bytes per guarded request | 32 MiB |
| Accumulated event payload bytes per guarded request | 32 MiB |
| Accumulated buffered events per guarded request | 65,536 |
| Private continuation bytes per guarded request, including tool results | 32 MiB |
| Private continuation items per guarded request | 65,536 |
| Default timeout for each judge call | 10 minutes, configured by GUARDRAIL_JUDGE_TIMEOUT |
The byte and event budgets span all model turns in the request. Releasing a turn's events does not reset them. The private-context budget is independent of the event budget and is released at request end. These limits account for raw bytes, not total heap use; parsed objects and temporary serialization need additional memory. If a limit is exceeded, the request fails without releasing its protected answer. Streaming clients should allow for the time spent generating and judging the answer. Higher reasoning effort can increase that wait; low is the default. Four sequential checks can consume up to forty minutes of judge time at the default timeout, but a block, error, cancellation, or another request timeout can end the request sooner.
Transcribing audio
POST /v1/audio/transcriptions turns an audio file into text. Send it as a multipart form:
curl -sS "https://api.<your-domain>/v1/audio/transcriptions" \
-H "Authorization: Bearer cm_api_…" \
-F "model=YOUR_TRANSCRIPTION_MODEL_ID" \
-F "[email protected]"
model takes the id of a transcription model. GET /v1/models lists chat models by default, so ask for these with ?type=transcription.
The platform includes a transcription model you can deploy. Select the Whisper Large v3 preset on the portal's Models page. See Deploying a model.
Any other field — language, prompt, response_format, timestamp_granularities[] — is passed to the model unchanged, and its answer is returned unchanged, so what you can send and what you get back depend on the model you choose. Two fields are the exception and never reach the model: tenant_id, because your identity comes from your API key, and cache_salt, which the gateway sets itself rather than letting a request choose one.
Transcription requests can be up to 128 MiB, including the audio file, form fields, and multipart overhead. The Chat UI accepts audio files up to 127 MiB, leaving room for that overhead. Other inference endpoints retain their 32 MiB request cap. The selected backend may impose its own limits.
Streaming a transcription
Add -F "stream=true" to receive the text as Server-Sent Events instead of waiting for the whole file. Not every model can stream; when the model replies in one piece the gateway returns that single JSON answer, so handle both cases.
Errors
Check the HTTP status before reading the response body. Most API errors use this JSON shape —
{"error": {"message": …, "type": …}} — but clients should not require a JSON body to handle an error:
| Status | Meaning |
|---|---|
400 | Malformed or unsupported request |
401 or 403 | Authentication failed (e.g. disabled or revoked key), or missing model/alias invoke grant |
404 | The model doesn't exist or isn't ready |
413 | The request exceeds 128 MiB for audio transcriptions, or 32 MiB for other inference endpoints |
501 | The request uses a feature the Model Gateway does not implement |
502 | The backend failed or was unreachable |
Model backend errors on the Responses API
The following behavior applies when the selected model backend rejects a request to /v1/responses. It does not apply
to gateway errors that happen before the model is invoked, such as an invalid platform API key or missing invoke
permission.
The gateway does not return provider-owned error messages, types, or codes to the client. Clients receive image-specific guidance when the model service identifies unsupported image input. Unsupported audio, video, and unidentified input modalities use the generic invalid-request message.
The client receives a stable error using one of these public messages:
| Failure classification | Public message | Meaning |
|---|---|---|
| Unsupported input modality explicitly identified as image | The selected model does not support image input. Remove the image or choose a model that supports images, then try again. | The model service specifically identified image input as unsupported. |
| Unsupported audio, video, or unidentified input modality | The selected model could not process this request. Check your input and try again. | The model service rejected an input modality, but the gateway does not currently provide modality-specific guidance for it. |
Context-window limit, usually 400 or 422 | This request is too long for the selected model. Shorten the conversation or attached content, then try again. | The request exceeds the model's context window. |
401 from the model service | The model service could not authenticate the request. Contact your administrator. | The credentials held by the platform for the model service were rejected. This does not refer to the caller's platform API key. |
403 from the model service | The model service denied this request. Contact your administrator if you need access. | The model service refused the request. |
408 or 504 | The model took too long to respond. Try again. | The model did not respond before the timeout. |
429 | The model is receiving too many requests. Wait a moment and try again. | The model service is limiting the request rate. |
Any other 4xx error | The selected model could not process this request. Check your input and try again. | The model rejected the request for another reason. |
Any other 5xx error | The model service is temporarily unavailable. Try again. | The model service failed or the gateway could not reach it. |
When the backend rejection is returned as an ordinary HTTP response, the gateway preserves its 400–599 status. If
a direct 429 response includes a Retry-After header, the gateway forwards that header.
Once a Responses API response has begun, including during a gateway tool loop, the HTTP status can remain 200. For a
non-streaming request, inspect the response's status and error fields. For a streaming request, a model-backend
failure is returned as a canonical response.failed event. Its response.error object contains the public message,
type, and code. A normally terminated stream then ends with [DONE].
The gateway also fails the response when the model service returns an unsupported hosted-tool call, such as
web_search_call. The gateway does not run that tool call and does not send it to the client. For a streaming request, the client
receives a response.failed event whose response.error.message is backend returned unsupported tool call output,
followed by [DONE]. Text that arrived before the failure stays in the stream.
Provider details are not returned to the client. Ask your administrator to inspect the Model Gateway diagnostics when the exact provider reason is needed.
This normalization is specific to /v1/responses. It does not change the behavior of other model endpoints.
Model capabilities
Every model advertises what it can do, so callers and the UI can pick the right model for a task. Capabilities have three independent axes:
-
type— the task family the model serves. Exactly one applies:typeUsed for chatChat Completions and Responses requests embeddingEmbedding requests and RAG indexing rerankRelevance scoring inside deployed RAG services transcriptionAudio transcription requests ocrDocument text extraction through Chat Completions and deployed RAG services guardrailPolicy and safety evaluation More surfaces (image generation, …) are planned for a future release.
-
input_modalities— what the model accepts:text,image,audio,video. -
output_modalities— what the model produces:text,embedding,image,audio.
type describes the API surface, not every capability. A vision-capable chat model is still type: chat; its image support is expressed as input_modalities: [text, image]. Keeping the task type and the modalities as separate axes (rather than one flat list) means a model always has exactly one task, while still describing the data shapes it handles.
A model whose type was never set is treated as chat (backward-compatible default), so models that predate this feature keep working until they are re-typed.
Listing and filtering models
GET /v1/models returns the models the caller is allowed to invoke. The response is OpenAI-compatible, with capability fields added:
{
"object": "list",
"data": [
{
"id": "chat-model-id",
"object": "model",
"owned_by": "confidentialmind",
"display_name": "Chat model",
"type": "chat",
"input_modalities": ["text"],
"output_modalities": ["text"]
}
]
}
Default: chat only
By default GET /v1/models returns only chat models:
curl -sS "https://api.<your-domain>/v1/models" \
-H "Authorization: Bearer <token>"
This keeps chat-first clients clean: non-chat models do not clutter the model picker and appear only when asked for.
Guardrail models are therefore omitted from this default list and from chat model pickers. Use type=guardrail to list
models intended for policy and safety checks.
Filtering
Three query parameters filter the list, combined with AND:
type— one of the values above,agentfor preset agents (see below), orallto disable the type filter.input_modalities— repeatable and/or comma-separated.output_modalities— repeatable and/or comma-separated.
Within a modality parameter the model must support every value listed.
# Everything, regardless of type
curl -sS "https://api.<your-domain>/v1/models?type=all" -H "Authorization: Bearer <token>"
# Just embedding models (e.g. to configure a RAG embedder)
curl -sS "https://api.<your-domain>/v1/models?type=embedding" -H "Authorization: Bearer <token>"
# Just guardrail models
curl -sS "https://api.<your-domain>/v1/models?type=guardrail" -H "Authorization: Bearer <token>"
# Vision-capable chat models (accept image input, produce text)
curl -sS "https://api.<your-domain>/v1/models?type=chat&input_modalities=image&output_modalities=text" \
-H "Authorization: Bearer <token>"
Quote the URL — an unquoted
?/&is interpreted by your shell.
An unknown type or modality value returns 400 Bad Request, so typos surface instead of silently returning an empty list.
Preset agents
GET /v1/models?type=agent returns the preset agents the
caller is allowed to invoke. Agents appear only under type=agent, and type=agent returns nothing else.
They are not included in type=all or in any model type.
An agent entry always reads "type": "agent". Its input_modalities and output_modalities are those of the
model the agent runs on, so an agent on a vision model lists image input and accepts image_url content
parts like that model does. The modality filters work on agents too:
# Preset agents that accept image input
curl -sS "https://api.<your-domain>/v1/models?type=agent&input_modalities=image" \
-H "Authorization: Bearer <token>"
An agent whose model can no longer be found stays in the list and reads as text in, text out.
How a model's capabilities are set
- Catalog (preset) deployments are typed automatically from the preset (e.g. the embedding preset deploys an
embeddingmodel). No extra input needed. - Custom and external models default to
chat; choose the correct type in the model create form before deploying. - Existing deployments can be re-typed at any time from the model's Edit page — useful for classifying models created before capabilities existed (which otherwise read as
chat). Editing the type is metadata-only and does not redeploy the model.
Input/output modalities are derived from the type by default (e.g. embedding → output [embedding]); set them explicitly when a model is multimodal (e.g. add image input for a vision chat model).