Skip to main content

Responses

POST /v1/responses implements the OpenAI Responses API. The OpenAI SDKs work against it with a base_url change, and any model in the catalog can serve a Responses request, not just OpenAI's. This is also the format OpenAI Codex speaks.

Requests and responses use the Responses shapes throughout: SDK helpers like response.output_text and client.responses.stream() work, function-tool round trips work on any model, and turns can be chained by previous_response_id instead of resending the conversation. One gap worth knowing before you start: OpenAI's separate Conversations API (the conversation parameter) is not supported; it is rejected with a 400. Chain turns with previous_response_id instead.

Endpoint​

POST https://api.mindshub.ai/v1/responses, plus the stored-response operations GET/DELETE /v1/responses/{id}, POST /v1/responses/{id}/cancel, and GET /v1/responses/{id}/input_items.

OpenAI SDKs take the base URL with /v1 and authenticate with api_key.

Quick example​

from openai import OpenAI
import os

client = OpenAI(
base_url="https://api.mindshub.ai/v1",
api_key=os.environ["MINDSHUB_API_KEY"],
)

response = client.responses.create(
model="mindshub_air",
input="What is the capital of Australia?",
)
print(response.output_text)

Request parameters​

ParameterHandlingNotes
modeladaptedAn alias from GET /v1/models resolves to a concrete provider model, so any catalog model can serve a Responses request. The response names the model that served, as every lane does; mindshub_air and mindshub_blaze return their own alias instead.
inputadaptedA string or an item list. message, function_call, function_call_output, custom_tool_call, custom_tool_call_output and reasoning items round-trip; input_text/output_text/input_image content parts translate, including inside a function_call_output, so a tool that returns an image shows the model an image; a model that cannot take one sees a short text note in its place. Supplied refusal explanations replay as text. input_file and input_audio parts are dropped (Milestone 4). An item_reference item has no local counterpart and is skipped. Every skipped item type (item_reference, local_shell_call, …) is named in X-MindsHub-Dropped-Params as input.<type>.
instructionsadaptedBecomes a leading system message. Deliberately per-turn: OpenAI does not carry it across previous_response_id, so a chained turn replays history without it.
streamforwardedSSE Responses events terminated by response.completed/.incomplete, never a [DONE] sentinel. Items are strictly sequential — one closes before the next opens.
stream_optionsignoredThe Responses API's only member is include_obfuscation, which we do not implement; usage rides the terminal event unconditionally, so there is nothing here to honour. include_usage is read as a MindsHub extension for callers migrating from the chat lane, and changes nothing a Responses client can observe.
max_output_tokensconditionalClamped down to the resolved model's output ceiling and up to the transport's floor; either is named in X-MindsHub-Clamped-Params. A turn cut short by it reports status: incomplete with incomplete_details.reason: max_output_tokens.
temperatureconditionalDropped on models that reject sampling params and named in X-MindsHub-Dropped-Params. One body routes to seven providers, so a param the chosen model cannot read is not a reason to refuse the request.
top_pconditionalAs temperature. On a Claude model that still takes sampling params, sending it with temperature drops top_p (Anthropic rejects the pair) and names it in X-MindsHub-Dropped-Params; top_p alone is forwarded.
verbosityconditionalNot in the SDK's request-params object. responses.parse() takes it as a top-level keyword and passes it through verbatim, so it arrives beside text rather than inside it; responses.create() does not model it at all, so parse() is the only SDK route to this parameter. Normalized into text.verbosity on arrival, so the echo and the applied value agree, and honoured from there.
reasoningconditionaleffort maps onto each model's own ladder, clamped rather than dropped so a request to be cheap is not answered by the provider's near-the-top default. summary is surfaced on the returned reasoning item where the provider gives one. Kimi reasons internally with no adjustable level and declares effort unsupported.
includeconditionalOf the eight values OpenAI defines, only reasoning.encrypted_content is honoured — it gates the reasoning round-trip. Note that web_search_call items and output_text annotations are returned whether or not they are asked for, so the two include values naming them change nothing. message.output_text.logprobs needs logprobs we do not produce; the file_search, code_interpreter and computer_use values belong to hosted tools we do not run; message.input_image.image_url is Milestone 4.
toolsadaptedResponses function tools are flat where the internal shape nests them; both translate. Hosted web_search maps onto each provider's own search or our external loop, and the searches it runs are reported as web_search_call output items on a non-streamed turn (a stream does not narrate them yet). mcp runs on the target's own remote-MCP connector (GPT: OpenAI's, require_approval passed through, approvals included; Claude: Anthropic's, which has no gate, so only tools resolving to never run and the rest are named in X-MindsHub-Dropped-Params) and is reported as mcp_list_tools / mcp_call output items on a non-streamed turn. A freeform custom tool (Codex's apply_patch) goes upstream as a function taking one string, its grammar in the description, and the call comes back as a custom_tool_call. Every other hosted tool (file_search, code_interpreter, computer_use, local_shell, tool_search, …) has no cross-provider meaning, so it is dropped and named in X-MindsHub-Dropped-Params as tools.<type>.
tool_choiceadaptedauto/required/none and {type: function, name} all map. A forced choice a model cannot honour is downgraded to auto rather than refused (ENG-1095). A hosted-tool choice cannot be forced generically and is dropped.
parallel_tool_callsconditionalHonoured, including on the Anthropic wire, which spells it inverted inside tool_choice. Gemini cannot express it and declares it unsupported, so the drop is named in X-MindsHub-Dropped-Params.
textconditionaltext.format is structured output and is honoured on every transport that can express a JSON schema; a model that cannot is a 400 naming it, not a silent drop, because prose handed to a caller about to json.loads() it is a different kind of answer. Schema-less json_object is unavailable on the Anthropic family, which has no equivalent. text.verbosity reaches only the OpenAI Responses transport, whose native field it is.
storeadaptedDefaults to true as OpenAI's server does. A stored turn is readable by GET, DELETE and /input_items, and chainable by previous_response_id, for 30 days. store: false persists nothing — a chain can still be read from, but not extended.
previous_response_idadaptedReconstructs history from the stored chain. Items the client also re-sent are deduped, because a duplicated tool_use is a hard 400 on the Anthropic wire. An unknown id is a 404 naming the recovery. A chain truncated to fit is reported in a response header.
backgroundignoredAccepted and echoed, but every request is served synchronously: a client's poll of retrieve() finds the response already completed rather than queued. response.queued is therefore never emitted, and cancel() has nothing in flight to stop.
metadataignoredAccepted and echoed back on the Response. Describes the caller's bookkeeping rather than the generation, so it reaches no provider and is not reported as a drop.
userignoredAs safety_identifier — OpenAI's deprecated spelling of it.
safety_identifierignoredAccepted and echoed. Describes the caller, not the generation. Abuse attribution here runs off the authenticated identity, which a request cannot choose.
service_tierignoredAccepted and echoed as default. One serving tier here, so the field is a constant rather than a measurement — sent because OpenAI always sends it and a client reading it should get a string rather than a missing key.
prompt_cache_keyignoredAccepted and echoed. We derive an equivalent key ourselves from the instructions and tool set, so honouring the caller's spelling would change nothing they can observe.
prompt_cache_retentionignoredAccepted and echoed. Cache lifetime is the upstream provider's to set and differs per provider; there is no field to forward it to on any transport we speak.
top_logprobsignoredOur OpenAI and xAI providers speak the Responses API, which has no logprobs, and the Anthropic wire has none either — so there is no response-side plumbing to build on for the two backends that could express it. Named in X-MindsHub-Dropped-Params, and logprobs: [] is sent on every text frame because the SDK requires the key.
truncationignoreddisabled IS our behavior and is honoured by doing nothing. auto asks us to drop history to fit the context window; we have no per-model context table, so it is named in X-MindsHub-Dropped-Params rather than silently not happening.
max_tool_callsignoredNo transport we speak has a per-turn tool-call ceiling, and enforcing one ourselves would mean truncating a turn mid-flight. Named in X-MindsHub-Dropped-Params.
conversationrejected (400)A 400. OpenAI documents conversation and previous_response_id as mutually exclusive, so the alternative is a one-line change for the caller. Refusing beats ignoring here: a client that believes we are holding its thread stops sending history, and every turn silently loses its context (ENG-1224). No third-party Responses implementation offers a Conversations API today.
promptrejected (400)A 400. Stored prompt templates live in OpenAI's dashboard and we cannot resolve an id that only exists there; serving the request would answer with the template silently missing.
context_managementrejected (400)A 400. It asks us to manage the context window on the caller's behalf, and a client that believes we are doing so stops managing it itself — the same failure as conversation.

This table is generated from PARAM_SUPPORT in minds/requests/responses_compat.py, which a unit test holds to the vendor SDK's own parameter list. The cross-API view is on the capability matrix.

conversation, prompt and context_management are rejected with a 400 naming the field rather than silently ignored: a client that believes we are holding its thread would otherwise stop sending history and lose context on every turn.

The cross-API view of every parameter is on the capability matrix.

Input​

input takes a plain string, or an array of items for multi-turn conversations and richer content:

{
"model": "sonnet",
"input": [
{"role": "user", "content": "What's in this image?"},
{"role": "user", "content": [
{"type": "input_text", "text": "Describe the chart."},
{"type": "input_image", "image_url": "https://example.com/chart.png"}
]}
]
}

Supported item content types are input_text, output_text, and input_image. Tool round-trips use function_call and function_call_output items; see Tool calling. Roles are system, user, and assistant. Prefer instructions over a system item for system-level guidance; both work.

Request body​

The full request and response schemas, with every field described, are on the generated reference pages: POST /v1/responses, GET /v1/responses/{id}, DELETE /v1/responses/{id}, cancel, and input_items.

Response​

{
"id": "resp_9f2c1ae0b4d8",
"object": "response",
"created_at": 1785401283,
"status": "completed",
"model": "claude-sonnet-5",
"output": [
{
"type": "message",
"id": "msg_1c40a2",
"role": "assistant",
"status": "completed",
"content": [
{"type": "output_text", "text": "Canberra.", "annotations": []}
]
}
],
"output_text": "Canberra.",
"usage": {
"input_tokens": 12,
"input_tokens_details": {"cached_tokens": 0},
"output_tokens": 4,
"output_tokens_details": {"reasoning_tokens": 0},
"total_tokens": 16
}
}

Notes on the shape:

  • output_text is the flattened assistant text, the same convenience field the OpenAI SDK exposes. output carries the structured items: message items, function_call items when the model calls tools, web_search_call items when it searched, mcp_list_tools / mcp_call / mcp_approval_request items when it used a remote MCP server, and a reasoning item on reasoning models.
  • model names the model that served, as on Chat Completions; mindshub_air and mindshub_blaze return their own alias. See Models.
  • input_tokens_details.cached_tokens reports prompt cache reads. See Prompt caching.
  • output_tokens_details.reasoning_tokens is a subset of output_tokens, broken out where the provider reports it. See Reasoning.
  • status is completed, incomplete (with incomplete_details.reason set to max_output_tokens or content_filter), or failed.
  • When the provider explains a refusal, the response includes a {"type": "refusal", "refusal": "..."} content part carrying that explanation, in place of (or beside) the output_text part. No refusal text is invented when the provider supplies none. output is never empty: a turn that produced nothing still returns a message item with empty text, so "declined" and "said nothing" stay distinguishable.
  • output_text.annotations carries the model's web-search sources as url_citation entries. See Web search.

Streaming events​

"stream": true produces the typed Responses event stream (response.created … response.completed), 15 of the protocol's 53 event types. The sequence, the terminal response.incomplete/response.failed events, and the list of events this API never emits are in Streaming.

Lane-specific behaviour​

  • Server-side state. Turns are stored for 30 days by default and chain by previous_response_id; GET, DELETE, input_items, and cancel operate on stored turns. The boundaries (25-turn chain walk, X-MindsHub-Chain-Truncated, expiry) are in Conversation state.
  • background: true is accepted but every request completes synchronously; cancel on anything else returns 400.
  • Trailing slash. POST /v1/responses/ is served identically to POST /v1/responses, so a client that appends the slash is not redirected with a body-dropping 307.
  • Codex speaks this format exclusively, so wire_api = "responses" runs it on any catalog model. See Codex.

Errors​

Errors use the OpenAI envelope, so SDK error handling works unmodified:

{
"error": {
"message": "The model 'foo' does not exist or you do not have access to it.",
"type": "invalid_request_error",
"param": "model",
"code": "model_not_found"
}
}

A rejected conversation, prompt, or context_management field is 400 unsupported_parameter; an unknown previous_response_id is 404 previous_response_not_found. Status codes and their meanings are shared across the API; see Errors.

The per-parameter contract this page summarizes is maintained in code, as PARAM_SUPPORT in minds/requests/responses_compat.py, and held to the OpenAI SDK's own parameter list by a test.