Responses
POST /v1/responses implements the OpenAI Responses API. The OpenAI SDKs work against it with a base_url change, and any model in the catalog can serve a Responses request, not just OpenAI's. This is also the format OpenAI Codex speaks.
Requests and responses use the Responses shapes throughout: SDK helpers like response.output_text and client.responses.stream() work, function-tool round trips work on any model, and turns can be chained by previous_response_id instead of resending the conversation. One gap worth knowing before you start: OpenAI's separate Conversations API (the conversation parameter) is not supported; it is rejected with a 400. Chain turns with previous_response_id instead.
Endpoint
POST https://api.mindshub.ai/v1/responses, plus the stored-response operations GET/DELETE /v1/responses/{id}, POST /v1/responses/{id}/cancel, and GET /v1/responses/{id}/input_items.
OpenAI SDKs take the base URL with /v1 and authenticate with api_key.
Quick example
- Python
- TypeScript
from openai import OpenAI
import os
client = OpenAI(
base_url="https://api.mindshub.ai/v1",
api_key=os.environ["MINDSHUB_API_KEY"],
)
response = client.responses.create(
model="mindshub_air",
input="What is the capital of Australia?",
)
print(response.output_text)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.mindshub.ai/v1",
apiKey: process.env.MINDSHUB_API_KEY,
});
const response = await client.responses.create({
model: "mindshub_air",
input: "What is the capital of Australia?",
});
console.log(response.output_text);
Request parameters
| Parameter | Handling | Notes |
|---|---|---|
model | adapted | An alias from GET /v1/models resolves to a concrete provider model, so any catalog model can serve a Responses request. The response names the model that served, as every lane does; mindshub_air and mindshub_blaze return their own alias instead. |
input | adapted | A string or an item list. message, function_call, function_call_output, custom_tool_call, custom_tool_call_output and reasoning items round-trip; input_text/output_text/input_image content parts translate, including inside a function_call_output, so a tool that returns an image shows the model an image; a model that cannot take one sees a short text note in its place. Supplied refusal explanations replay as text. input_file and input_audio parts are dropped (Milestone 4). An item_reference item has no local counterpart and is skipped. Every skipped item type (item_reference, local_shell_call, …) is named in X-MindsHub-Dropped-Params as input.<type>. |
instructions | adapted | Becomes a leading system message. Deliberately per-turn: OpenAI does not carry it across previous_response_id, so a chained turn replays history without it. |
stream | forwarded | SSE Responses events terminated by response.completed/.incomplete, never a [DONE] sentinel. Items are strictly sequential — one closes before the next opens. |
stream_options | ignored | The Responses API's only member is include_obfuscation, which we do not implement; usage rides the terminal event unconditionally, so there is nothing here to honour. include_usage is read as a MindsHub extension for callers migrating from the chat lane, and changes nothing a Responses client can observe. |
max_output_tokens | conditional | Clamped down to the resolved model's output ceiling and up to the transport's floor; either is named in X-MindsHub-Clamped-Params. A turn cut short by it reports status: incomplete with incomplete_details.reason: max_output_tokens. |
temperature | conditional | Dropped on models that reject sampling params and named in X-MindsHub-Dropped-Params. One body routes to seven providers, so a param the chosen model cannot read is not a reason to refuse the request. |
top_p | conditional | As temperature. On a Claude model that still takes sampling params, sending it with temperature drops top_p (Anthropic rejects the pair) and names it in X-MindsHub-Dropped-Params; top_p alone is forwarded. |
verbosity | conditional | Not in the SDK's request-params object. responses.parse() takes it as a top-level keyword and passes it through verbatim, so it arrives beside text rather than inside it; responses.create() does not model it at all, so parse() is the only SDK route to this parameter. Normalized into text.verbosity on arrival, so the echo and the applied value agree, and honoured from there. |
reasoning | conditional | effort maps onto each model's own ladder, clamped rather than dropped so a request to be cheap is not answered by the provider's near-the-top default. summary is surfaced on the returned reasoning item where the provider gives one. Kimi reasons internally with no adjustable level and declares effort unsupported. |
include | conditional | Of the eight values OpenAI defines, only reasoning.encrypted_content is honoured — it gates the reasoning round-trip. Note that web_search_call items and output_text annotations are returned whether or not they are asked for, so the two include values naming them change nothing. message.output_text.logprobs needs logprobs we do not produce; the file_search, code_interpreter and computer_use values belong to hosted tools we do not run; message.input_image.image_url is Milestone 4. |
tools | adapted | Responses function tools are flat where the internal shape nests them; both translate. Hosted web_search maps onto each provider's own search or our external loop, and the searches it runs are reported as web_search_call output items on a non-streamed turn (a stream does not narrate them yet). mcp runs on the target's own remote-MCP connector (GPT: OpenAI's, require_approval passed through, approvals included; Claude: Anthropic's, which has no gate, so only tools resolving to never run and the rest are named in X-MindsHub-Dropped-Params) and is reported as mcp_list_tools / mcp_call output items on a non-streamed turn. A freeform custom tool (Codex's apply_patch) goes upstream as a function taking one string, its grammar in the description, and the call comes back as a custom_tool_call. Every other hosted tool (file_search, code_interpreter, computer_use, local_shell, tool_search, …) has no cross-provider meaning, so it is dropped and named in X-MindsHub-Dropped-Params as tools.<type>. |
tool_choice | adapted | auto/required/none and {type: function, name} all map. A forced choice a model cannot honour is downgraded to auto rather than refused (ENG-1095). A hosted-tool choice cannot be forced generically and is dropped. |
parallel_tool_calls | conditional | Honoured, including on the Anthropic wire, which spells it inverted inside tool_choice. Gemini cannot express it and declares it unsupported, so the drop is named in X-MindsHub-Dropped-Params. |
text | conditional | text.format is structured output and is honoured on every transport that can express a JSON schema; a model that cannot is a 400 naming it, not a silent drop, because prose handed to a caller about to json.loads() it is a different kind of answer. Schema-less json_object is unavailable on the Anthropic family, which has no equivalent. text.verbosity reaches only the OpenAI Responses transport, whose native field it is. |
store | adapted | Defaults to true as OpenAI's server does. A stored turn is readable by GET, DELETE and /input_items, and chainable by previous_response_id, for 30 days. store: false persists nothing — a chain can still be read from, but not extended. |
previous_response_id | adapted | Reconstructs history from the stored chain. Items the client also re-sent are deduped, because a duplicated tool_use is a hard 400 on the Anthropic wire. An unknown id is a 404 naming the recovery. A chain truncated to fit is reported in a response header. |
background | ignored | Accepted and echoed, but every request is served synchronously: a client's poll of retrieve() finds the response already completed rather than queued. response.queued is therefore never emitted, and cancel() has nothing in flight to stop. |
metadata | ignored | Accepted and echoed back on the Response. Describes the caller's bookkeeping rather than the generation, so it reaches no provider and is not reported as a drop. |
user | ignored | As safety_identifier — OpenAI's deprecated spelling of it. |
safety_identifier | ignored | Accepted and echoed. Describes the caller, not the generation. Abuse attribution here runs off the authenticated identity, which a request cannot choose. |
service_tier | ignored | Accepted and echoed as default. One serving tier here, so the field is a constant rather than a measurement — sent because OpenAI always sends it and a client reading it should get a string rather than a missing key. |
prompt_cache_key | ignored | Accepted and echoed. We derive an equivalent key ourselves from the instructions and tool set, so honouring the caller's spelling would change nothing they can observe. |
prompt_cache_retention | ignored | Accepted and echoed. Cache lifetime is the upstream provider's to set and differs per provider; there is no field to forward it to on any transport we speak. |
top_logprobs | ignored | Our OpenAI and xAI providers speak the Responses API, which has no logprobs, and the Anthropic wire has none either — so there is no response-side plumbing to build on for the two backends that could express it. Named in X-MindsHub-Dropped-Params, and logprobs: [] is sent on every text frame because the SDK requires the key. |
truncation | ignored | disabled IS our behavior and is honoured by doing nothing. auto asks us to drop history to fit the context window; we have no per-model context table, so it is named in X-MindsHub-Dropped-Params rather than silently not happening. |
max_tool_calls | ignored | No transport we speak has a per-turn tool-call ceiling, and enforcing one ourselves would mean truncating a turn mid-flight. Named in X-MindsHub-Dropped-Params. |
conversation | rejected (400) | A 400. OpenAI documents conversation and previous_response_id as mutually exclusive, so the alternative is a one-line change for the caller. Refusing beats ignoring here: a client that believes we are holding its thread stops sending history, and every turn silently loses its context (ENG-1224). No third-party Responses implementation offers a Conversations API today. |
prompt | rejected (400) | A 400. Stored prompt templates live in OpenAI's dashboard and we cannot resolve an id that only exists there; serving the request would answer with the template silently missing. |
context_management | rejected (400) | A 400. It asks us to manage the context window on the caller's behalf, and a client that believes we are doing so stops managing it itself — the same failure as conversation. |
This table is generated from PARAM_SUPPORT in minds/requests/responses_compat.py, which a unit test holds to the vendor SDK's own parameter list. The cross-API view is on the capability matrix.
conversation, prompt and context_management are rejected with a 400 naming the field rather than silently ignored: a client that believes we are holding its thread would otherwise stop sending history and lose context on every turn.
The cross-API view of every parameter is on the capability matrix.
Input
input takes a plain string, or an array of items for multi-turn conversations and richer content:
{
"model": "sonnet",
"input": [
{"role": "user", "content": "What's in this image?"},
{"role": "user", "content": [
{"type": "input_text", "text": "Describe the chart."},
{"type": "input_image", "image_url": "https://example.com/chart.png"}
]}
]
}
Supported item content types are input_text, output_text, and input_image. Tool round-trips use function_call and function_call_output items; see Tool calling. Roles are system, user, and assistant. Prefer instructions over a system item for system-level guidance; both work.
Request body
The full request and response schemas, with every field described, are on the generated reference pages: POST /v1/responses, GET /v1/responses/{id}, DELETE /v1/responses/{id}, cancel, and input_items.
Response
{
"id": "resp_9f2c1ae0b4d8",
"object": "response",
"created_at": 1785401283,
"status": "completed",
"model": "claude-sonnet-5",
"output": [
{
"type": "message",
"id": "msg_1c40a2",
"role": "assistant",
"status": "completed",
"content": [
{"type": "output_text", "text": "Canberra.", "annotations": []}
]
}
],
"output_text": "Canberra.",
"usage": {
"input_tokens": 12,
"input_tokens_details": {"cached_tokens": 0},
"output_tokens": 4,
"output_tokens_details": {"reasoning_tokens": 0},
"total_tokens": 16
}
}
Notes on the shape:
output_textis the flattened assistant text, the same convenience field the OpenAI SDK exposes.outputcarries the structured items:messageitems,function_callitems when the model calls tools,web_search_callitems when it searched,mcp_list_tools/mcp_call/mcp_approval_requestitems when it used a remote MCP server, and areasoningitem on reasoning models.modelnames the model that served, as on Chat Completions;mindshub_airandmindshub_blazereturn their own alias. See Models.input_tokens_details.cached_tokensreports prompt cache reads. See Prompt caching.output_tokens_details.reasoning_tokensis a subset ofoutput_tokens, broken out where the provider reports it. See Reasoning.statusiscompleted,incomplete(withincomplete_details.reasonset tomax_output_tokensorcontent_filter), orfailed.- When the provider explains a refusal, the response includes a
{"type": "refusal", "refusal": "..."}content part carrying that explanation, in place of (or beside) theoutput_textpart. No refusal text is invented when the provider supplies none.outputis never empty: a turn that produced nothing still returns amessageitem with empty text, so "declined" and "said nothing" stay distinguishable. output_text.annotationscarries the model's web-search sources asurl_citationentries. See Web search.
Streaming events
"stream": true produces the typed Responses event stream (response.created … response.completed), 15 of the protocol's 53 event types. The sequence, the terminal response.incomplete/response.failed events, and the list of events this API never emits are in Streaming.
Lane-specific behaviour
- Server-side state. Turns are stored for 30 days by default and chain by
previous_response_id;GET,DELETE,input_items, andcanceloperate on stored turns. The boundaries (25-turn chain walk,X-MindsHub-Chain-Truncated, expiry) are in Conversation state. background: trueis accepted but every request completes synchronously;cancelon anything else returns400.- Trailing slash.
POST /v1/responses/is served identically toPOST /v1/responses, so a client that appends the slash is not redirected with a body-dropping 307. - Codex speaks this format exclusively, so
wire_api = "responses"runs it on any catalog model. See Codex.
Errors
Errors use the OpenAI envelope, so SDK error handling works unmodified:
{
"error": {
"message": "The model 'foo' does not exist or you do not have access to it.",
"type": "invalid_request_error",
"param": "model",
"code": "model_not_found"
}
}
A rejected conversation, prompt, or context_management field is 400 unsupported_parameter; an unknown previous_response_id is 404 previous_response_not_found. Status codes and their meanings are shared across the API; see Errors.
The per-parameter contract this page summarizes is maintained in code, as PARAM_SUPPORT in minds/requests/responses_compat.py, and held to the OpenAI SDK's own parameter list by a test.