Responses API
Use POST /v1/responses for text generation, streaming, conversations, and model tool use. The API follows the
OpenAI Responses request and response shapes where documented here, with platform-specific conversation handling and
gateway-hosted tools.
For a single prompt and answer, send a string in input. For multiple turns, either send the history as an input
array or let the platform store it and continue with conversation_id.
Before you begin
You need:
- the API URL shown in the platform portal
- an API key with
invokeaccess to a chat model - the model ID from
GET /v1/models
The raw HTTP examples use curl. Set these values once:
export API_BASE_URL="https://api.<your-domain>"
export API_KEY="cm_api_…"
export MODEL_ID="YOUR_MODEL_ID"
The Python examples use the OpenAI Python package and the same environment variables. Install the package with your project's usual Python dependency manager.
Generate a response
Send a non-streaming request when you want one JSON result:
curl -sS "${API_BASE_URL}/v1/responses" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
--data-binary @- <<JSON
{
"model": "${MODEL_ID}",
"input": "Explain confidential computing in two sentences.",
"store": false
}
JSON
The answer is in the output array. A message can contain one or more output_text parts:
{
"status": "completed",
"output": [
{
"type": "message",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "Confidential computing protects data while it is being processed..."
}
]
}
]
}
If the model calls a function that is undeclared or disallowed for the request, the output array includes a
function_call and a function_call_output with an error message. The gateway does not execute the function and lets
the model continue, so the response can still end with status completed. This also applies when no tools are declared.
Read the answer from the message items, as the examples on this page do.
See Tool access is decided for each request for details.
With the OpenAI Python client:
import os
from openai import OpenAI
client = OpenAI(
base_url=f"{os.environ['API_BASE_URL']}/v1",
api_key=os.environ["API_KEY"],
)
response = client.responses.create(
model=os.environ["MODEL_ID"],
input="Explain confidential computing in two sentences.",
store=False,
)
if response.status != "completed":
raise RuntimeError(f"Response ended with status {response.status}: {response.error}")
print(response.output_text)
For direct model requests, store: false makes this a one-shot request. Storage otherwise defaults to true, and the
gateway creates or continues a conversation independently of the selected model backend.
Give the model instructions
Use instructions for behavior that should apply to the current request:
{
"model": "YOUR_MODEL_ID",
"instructions": "Answer for a non-technical reader. Be concise.",
"input": "What is a trusted execution environment?",
"store": false
}
For direct model requests, parameters such as max_output_tokens, temperature, and top_p can be sent when the
selected model and its configuration support them. Preset agents use their configured parameters; callers cannot
override them.
Chat mode
Set chat_mode to true to use the instruction bundle that Chat requests:
{
"model": "YOUR_MODEL_ID",
"input": "Explain this formula.",
"chat_mode": true,
"store": false
}
This gateway extension works with direct models and preset agents on POST /v1/responses.
Omitting it or sending false disables the bundle. null and other JSON types return HTTP 400.
Chat sends true on every request, including follow-ups and retries. API callers can opt in too;
the flag selects behavior and does not establish client identity or grant permissions.
- Direct models: fixed Chat formatting guidance, tenant assistant instructions, any caller
instructions, then the current UTC date. - Preset agents: configured agent instructions, fixed Chat formatting guidance, then the current UTC date. Tenant assistant instructions do not apply, even when the agent's configured instructions are empty.
The gateway adds the currently callable tool names after these instructions on each model dispatch, including dispatches after tool results, independently of Chat mode. This context reflects the request, agent configuration, and existing permissions; it does not grant access to tools. Agent restrictions and configured output guardrails also apply independently. There is no separate UTC option.
The date is captured once per request and stays the same across tool turns, even if they cross UTC midnight. For direct models, tenant instructions are also loaded once per request. Later requests see updated tenant settings. If that lookup fails, the request returns an error before model execution. Agent requests do not perform the lookup.
The gateway consumes chat_mode before sending the request to the model provider. Injected instructions are not
saved as conversation messages or included in public instruction snapshots. Direct-model snapshots retain only
caller instructions; agent snapshots contain instructions: null. A model can still mention an injected fact in
its answer. Send chat_mode on each request; stored conversations do not remember the option.
With the Python SDK, pass the extension through extra_body:
response = client.responses.create(
model=os.environ["MODEL_ID"],
input="Explain this formula.",
store=False,
extra_body={"chat_mode": True},
)
Continue a conversation
You can keep history in your application or let the platform store it.
| Approach | When to use it |
|---|---|
Send an input array | Portable client-managed pattern; your application owns and resends the history. |
Use conversation_id | Platform-stored history; later requests send only the new input. |
Preset-agent requests accept only user messages from
the caller, so use platform-stored history instead of resending client-managed history. Storage is disabled by default:
set store: true on the first turn, then continue with the returned conversation_id. Do not combine
conversation_id with store: false; the request returns 400.
Send history yourself
Start with a user message, then add the response output and the next user message to the next request:
history = [{"role": "user", "content": "My deployment window starts at 09:00 UTC."}]
first = client.responses.create(
model=os.environ["MODEL_ID"],
input=history,
store=False,
)
history.extend(item.model_dump() for item in first.output)
history.append({"role": "user", "content": "What time does my deployment window start?"})
second = client.responses.create(
model=os.environ["MODEL_ID"],
input=history,
store=False,
)
print(second.output_text)
Set store: false on every client-managed turn. Storage is enabled by default, so omitting it would create a new
stored conversation for each request even though your application is already supplying the history.
Let the platform store history
On the first request, omit conversation_id. The response contains a new ID in conversation.id:
curl -sS "${API_BASE_URL}/v1/responses" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
--data-binary @- <<JSON
{
"model": "${MODEL_ID}",
"input": "My deployment window starts at 09:00 UTC. Remember that.",
"store": true
}
JSON
Copy the ID from the response's conversation object:
{
"conversation": {
"id": "YOUR_CONVERSATION_ID"
}
}
Pass that ID as conversation_id on the next request. You send only the new input; the gateway loads the stored
history and supplies it to the model:
curl -sS "${API_BASE_URL}/v1/responses" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
--data-binary @- <<JSON
{
"model": "${MODEL_ID}",
"conversation_id": "YOUR_CONVERSATION_ID",
"input": "What time does my deployment window start?"
}
JSON
The same flow with the OpenAI Python client uses extra_body because conversation_id is a platform extension:
first = client.responses.create(
model=os.environ["MODEL_ID"],
input="My deployment window starts at 09:00 UTC. Remember that.",
store=True,
)
if first.conversation is None:
raise RuntimeError("The gateway did not create a conversation")
conversation_id = first.conversation.id
second = client.responses.create(
model=os.environ["MODEL_ID"],
input="What time does my deployment window start?",
extra_body={"conversation_id": conversation_id},
)
print(second.output_text)
Keep the conversation_id in your application's session or database if the conversation must survive a process
restart. Start a new conversation by omitting it from a later request. Use the same API key for later turns.
Conversation behavior and limits
| Behavior | What to expect |
|---|---|
| Create | Omit conversation_id and set store: true. The response contains conversation.id. |
| Continue | Send the stored ID as conversation_id; the gateway supplies recent history before the new input. |
| History window | Up to 128 recent stored items are supplied. The model's context limit still applies. |
| Access | A conversation is scoped to its creator and tenant. Keep using the same API key. |
| Disable storage | Set store: false. No conversation is created and conversation is absent from the response. |
| Change models | Prefer the same model. Model-specific reasoning and tool items may be omitted after a change. |
| Manage conversations | The public /v1 API currently does not expose list, retrieve, or delete operations. |
When assistant text is saved before a response fails, the terminal event includes
response.conversation.output_item_id, the stored assistant item's ID. That partial answer remains in conversation
history and is supplied to the model on the next turn, subject to the history window above, unless it is deleted.
Clients with conversation management access can delete it through the separate ConnectRPC endpoint
POST /cmind.conversations.v1.ConversationItemService/DeleteItem, with conversationId set to conversation.id,
itemId set to conversation.output_item_id, and tenantId set to the conversation's tenant ID in the JSON body.
This is not a /v1 operation. Clients using only /v1 can omit conversation_id to start a new conversation
without that saved partial answer.
conversation_id is the supported mechanism for continuing platform-stored history. The previous_response_id and
conversation request fields are not currently supported and do not continue a platform conversation.
Stream a response
Set stream: true to receive Server-Sent Events. Text arrives in response.output_text.delta events, followed by a
terminal response.completed, response.incomplete, or response.failed event and [DONE].
stream = client.responses.create(
model=os.environ["MODEL_ID"],
input="Write a four-line poem about private AI.",
stream=True,
store=False,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.failed":
raise RuntimeError(event.response.error)
print()
When storage is enabled, the terminal event's response.conversation.id contains the conversation ID to use for the
next request.
A preset agent with output guardrails holds its
final answer until the checks finish. The stream sends response.created and content-free response.in_progress
keepalives while waiting. Intermediate assistant text, reasoning, tool calls and tool results stay private.
These keepalives do not report generation or tool progress. Clients displaying text deltas receive no answer
text until generation and judging finish, so users see a wait followed by the complete result.
The selected final answer or block message then arrives as a group of SSE events without pacing, followed by the
terminal event and [DONE].
Use tools
The Responses API can run function tools published by the Model Gateway, including RAG and connected MCP tools. The gateway can execute those tools and feed their results back to the model within the same request.
→ Use Model Gateway tools with the Responses API
Tool availability depends on the selected model's function-calling support and configuration. OpenAI-hosted tool types such as web search, file search, code interpreter, computer use, and image generation are not provided by this endpoint, regardless of the selected model.
A model can still return an unsupported hosted-tool call, such as web_search_call. The gateway does not run it
and does not send it to you. The response fails instead. A streaming request ends with a response.failed event and [DONE].
Text that already arrived stays in the stream.
See Model backend errors on the Responses API.
tool_choice support depends on the request target:
- Direct in-platform and custom OpenAI-compatible models support
auto(the default),none, andallowed_toolswith modeauto. - Direct models provided by OpenAI or Azure OpenAI also support
required, a named function, andallowed_toolswith moderequired. - Preset-agent requests do not accept
tool_choice; the agent uses its configured tools.
A named function uses {"type": "function", "name": "your_function"}. To restrict the callable set, use:
{
"type": "allowed_tools",
"mode": "auto",
"tools": [{"type": "function", "name": "your_function"}]
}
Unsupported choices return 400. A named function must appear in the request's tools array; allowed_tools can
only select from functions declared there.
Handle unsuccessful responses
Before generation begins, authentication, authorization, validation, and backend connection failures use HTTP error
statuses. After a response has begun, the HTTP status can remain 200; inspect the response status and error, or
the terminal streaming event, to determine the outcome.
If the backend stream ends unexpectedly, the gateway sends response.failed followed by [DONE] while the client
connection remains writable. Usage from completed model turns is still reported.
See Model backend errors on the Responses API for the stable client-facing errors and retry guidance.
Compatibility notes
This endpoint is a supported subset of the OpenAI Responses API. In particular:
- only
POST /v1/responsesis exposed background(even whenfalse) andprevious_response_idare not currently supported; omit these fields- platform conversations use
conversation_id tool_choicesupport varies by model provider and request target; see tool-choice support- function and MCP tools are supported; OpenAI-hosted tool types are not
- preset-agent requests accept only user messages, use configured model parameters, and default to
store: false - preset agents with output guardrails delay the final answer until checks finish; see Stream a response
- input, output, tool use, sampling, and multimodal support can vary by model
Use this page as the Model Gateway contract. Fields and operations not documented here are not part of the supported public surface.