Skip to main content

Chat Completions

POST /v1/chat/completions follows the OpenAI chat completions shape: send a model and a list of messages, get back a completion. It is the most widely supported request format in the ecosystem and the one to start with. The OpenAI SDKs work against it with only a base_url change, and any model in the catalog can serve it.

Endpoint​

POST https://api.mindshub.ai/v1/chat/completions

OpenAI SDKs take the base URL with /v1 and authenticate with api_key. Install the official client: pip install openai or npm install openai.

Quick example​

from openai import OpenAI
import os

client = OpenAI(
base_url="https://api.mindshub.ai/v1",
api_key=os.environ["MINDSHUB_API_KEY"],
)

response = client.chat.completions.create(
model="mindshub_air",
messages=[
{"role": "system", "content": "You are a terse assistant."},
{"role": "user", "content": "What is the capital of Australia?"},
],
)
print(response.choices[0].message.content)

The samples call mindshub_air, the alias your recurring included allowance covers, so they work on an account with no payment method on file. Change the model string to any alias in Models once you have a wallet balance; nothing else about the request changes. The TypeScript sample uses top-level await, so run it as an ES module.

Request parameters​

ParameterHandlingNotes
modeladaptedAn alias from GET /v1/models resolves to a concrete provider model. The response names the model that served, which can be sent back as-is (it resolves to the alias serving that model, not to a pin); mindshub_air and mindshub_blaze return their own alias instead. The deprecated latest:<alias> spelling still resolves.
messagesadaptedsystem/user/assistant/tool all round-trip, and developer is read as system. Text and image_url content parts translate to every provider; file and input_audio parts do not (Milestone 4). name survives only to Moonshot, and as the tool-name fallback on Gemini. hosted_tool_calls is an internal carrier for the other two lanes' replayed remote-MCP records and is ignored here.
streamforwardedSSE chat.completion.chunk frames terminated by [DONE]. Null-valued keys are omitted rather than sent as null, so non-terminal chunks carry no finish_reason key.
stream_optionsadaptedinclude_usage is honoured; other keys are ignored. Never reaches a provider — it shapes our own response — so no model can fail to support it. Kimi streams end with a usage chunk whether or not it was asked for.
max_tokensconditionalClamped down to the resolved model's output ceiling and up to the transport's floor; either is named in X-MindsHub-Clamped-Params. Above the 131,072 admission ceiling the request is refused with max_tokens_exceeded before a provider is touched.
max_completion_tokensadaptedOpenAI's newer spelling of max_tokens and wins over it when both are set to positive values. A value of 0 is falsy here and falls through to max_tokens.
temperatureconditionalDropped on models that reject sampling params — every current Claude 5 model, Gemini 3.x, and the GPT 5.6/5.3-codex line — and named in X-MindsHub-Dropped-Params.
top_pconditionalAs temperature. On a Claude model that still takes sampling params, sending it with temperature drops top_p (Anthropic rejects the pair) and names it in X-MindsHub-Dropped-Params; top_p alone is forwarded.
top_kconditionalA MindsHub extension, not an OpenAI parameter. Forwarded where the transport has it; the OpenAI Responses shape does not, so it is always dropped on the GPT and Grok families and named in X-MindsHub-Dropped-Params.
stopconditionalForwarded where the transport has a stop-sequence field; the OpenAI Responses shape, which serves the GPT and Grok families here, does not. Reported under the internal name stop_sequences, not stop.
seedconditionalReaches Gemini and Moonshot, the only two transports with a field for it. Never a reproducibility guarantee even there.
presence_penaltyconditionalAs seed.
frequency_penaltyconditionalAs seed.
verbosityconditionalRides text.verbosity on the OpenAI Responses transport, the only one that has it; dropped and named everywhere else.
reasoning_effortconditionalClamped onto the resolved model's published ladder rather than dropped — with no level on the request the provider applies its own, which sits near the top. An unrecognized level fits no rung, so it falls back to the model's published default_reasoning_effort clamped onto that ladder, and is reported as a clamp naming the category other rather than the value sent; it is dropped only where the model publishes no default. A model that publishes a default and an empty ladder is pinned: every request runs at that level and a level sent is replaced by it (mindshub_air). Kimi reasons internally with no adjustable level, so the transport declares it unsupported.
toolsadaptedFunction tools translate to every provider, strict included. The platform's own web_search/fetch entries map onto each provider's hosted search, or our external loop, and are dropped where the target has neither. An OpenAI-shaped mcp entry (MindsHub extension: Chat Completions itself has none) runs on the target's own remote-MCP connector; this lane cannot hand back an approval request, so only tools with require_approval: "never" run and the rest are named in X-MindsHub-Dropped-Params. The answer reflects the tool output; the body has no slot for the calls.
tool_choiceadaptedauto/required/none/named all map. A forced choice a model cannot honour is downgraded to auto rather than refused (ENG-1095), and none is honoured by dropping the tools on transports with no word for it — both reported in X-MindsHub-Clamped-Params.
parallel_tool_callsconditionalHonoured, including on the Anthropic wire, which spells it inverted inside tool_choice. Gemini cannot express it and declares it unsupported, so the drop is named.
web_search_optionsadaptedRead as the platform's own {'type': 'web_search'} tool entry, so it enables hosted search on models that have it. search_context_size and user_location are dropped: one body here routes to seven backends whose search tools take different options.
functionsadaptedOpenAI's pre-tools spelling, deprecated by OpenAI and still served by them. Honoured in all four places they switch: the request becomes tools, a replayed assistant.function_call and 'function' role become tool_calls and a 'tool' message, and the answer comes back as message.function_call with finish_reason 'function_call' on both the JSON and SSE paths. Only the first call of a multi-call turn is representable. See minds/requests/legacy_functions.py.
function_calladaptedTranslated into tool_choice; an unrecognized string is dropped rather than forwarded into a provider 400. Sending tools alongside either legacy field is read as a modern request, and answered in the modern shape.
response_formatconditionaljson_schema and json_object both parse; each transport renders the schema in its own shape. The ONE param that raises rather than being dropped when a model cannot honour it (CONTRACT_PARAMS): prose handed to a caller who will json.loads() it is a different kind of answer, not a differently-flavored one. json_object is a 400 on the Anthropic wire, which has no schema-less JSON mode.
storeignoredChat completions are not stored and there is no retrieval endpoint for them. /v1/responses does store, and defaults to it.
metadataignoredAccepted and not read. Must be an object if present, so a string is a 400.
userignoredRequests are attributed by API key and the headers the edge supplies. Not reported as a drop: nothing about the answer changes.
safety_identifierignoredAs user.
service_tierignoredOne serving tier; nothing to select. The response echoes default.
prompt_cache_keyignoredCache routing is derived from the request's own stable prefix, which is more reliable than a key a caller has to keep consistent themselves.
prompt_cache_retentionignoredRetention is the upstream provider's.
nignoredThere is always exactly one choice. Reported in X-MindsHub-Dropped-Params unless it was 1, which is what we do.
logprobsignoredNo transport here returns log probabilities: they are Chat Completions fields, and the shape serving the GPT and Grok families is the Responses API. Reported as a drop unless explicitly false.
top_logprobsignoredAs logprobs.
logit_biasignoredAs logprobs. Reported as a drop.
predictionignoredPredicted outputs are not offered. Reported as a drop.
modalitiesignoredText only; audio is Milestone 4. Reported as a drop.
audioignoredAs modalities.

This table is generated from PARAM_SUPPORT in minds/requests/chat_completions_request.py, which a unit test holds to the vendor SDK's own parameter list. The cross-API view is on the capability matrix.

Every parameter OpenAI defines is accepted, and any other top-level parameter is accepted and ignored: unknown parameters are tolerated for SDK compatibility, since coding agents and SDKs treat request rejections as hard errors. Unknown nested content, like an unrecognized message part type, can reach the provider and fail there.

What the API can't honor, it names in X-MindsHub-Dropped-Params rather than discarding silently; the rows marked ignored above say which fields those are, and which are not reported because they describe your account rather than the generation. The one exception is a stream whose 200 went out early, which carries no X-MindsHub-* header; see Errors.

The cross-API view of every parameter is on the capability matrix.

Message content​

Roles are system, user, assistant, and tool. developer is accepted and read as system: it is OpenAI's newer spelling, recommended for reasoning models, and the two mean the same thing here.

content can be a plain string or an array of parts. Supported part types are text and image_url; see Images for the image rules.

Request body​

The full request and response schemas, with every field described, are on the generated reference page: POST /v1/chat/completions. The spec itself is at openapi.json.

Response​

{
"id": "chatcmpl-636a4f9b-80c4-4989-a48d-0c9bed215ba7",
"object": "chat.completion",
"created": 1785401283,
"model": "mindshub_air",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Canberra." },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 4,
"total_tokens": 16
}
}

Notes on the shape:

  • model names the model that served the request, fallback included, and can be sent again as-is. mindshub_air and mindshub_blaze return their own alias instead. See the models page.
  • There is always exactly one choice.
  • message.content is null (not an empty string) when the model produced no text, for example when it only called tools.
  • message.refusal carries the model's stated reason when it declined, on the models whose provider gives one (the Claude and GPT families). It is absent or null when the model refused without saying why. finish_reason is content_filter either way.
  • usage.completion_tokens_details is on every response. Its reasoning_tokens is a subset of completion_tokens, broken out by the GPT/Grok families and Gemini models; elsewhere it reads 0, which means no separately-reported reasoning rather than no reasoning. See Reasoning.
  • service_tier is always "default".
  • When the request touched a provider's prompt cache, usage gains prompt_tokens_details. See Prompt caching.
  • A turn that searched reports its sources on message.annotations. See Web search.

finish_reason​

ValueMeaning
stopThe model finished normally.
lengthOutput was truncated at max_tokens (or the model's context limit).
tool_callsThe model stopped to call one or more tools.
content_filterThe upstream provider refused to continue.
null / absentAbnormal end: the provider reported a failure or an unfinished turn. Treat as unsuccessful.

Model-family differences to know:

  • On the Claude family, mindshub_air, deepseek-v4-flash, qwen, glm, and muse-spark, a turn that was truncated and contains tool calls reports length (truncation wins). On Gemini models it reports tool_calls.
  • Gemini models never report null or content_filter: refusals and abnormal ends collapse into stop. Don't build safety-refusal detection on finish_reason for gemini/gemini-flash.
  • kimi responses come through with the provider's own finish reasons, unmapped.

Streaming events​

"stream": true returns text/event-stream: data: <json> lines carrying chat.completion.chunk frames, terminated by data: [DONE]. Usage arrives in a final choice-less chunk when you send stream_options: {"include_usage": true}. The frame format, the differences from OpenAI's stream, and the failure rules are in Streaming.

Lane-specific behaviour​

  • Funding is checked before the model runs. An empty wallet or exhausted allowance refuses the request up front (402 or 429) at no cost. See Billing.
  • Failover is reported. During a provider outage the platform may serve your request through a fallback model. X-MindsHub-Failover, X-MindsHub-Attempts and the body's mindshub field say so (a stream whose 200 went out early carries only the body field, and not even that if it fails), and X-MindsHub-Allow-Failover: false turns it off. See Failover.
  • Legacy functions/function_call are served in the legacy response shape. See Tool calling.

Errors​

Errors use the OpenAI envelope: {"error": {"message", "type", "param", "code"}}, so client.chat.completions.create(...) raises the SDK's typed errors. A 400 from a malformed request body carries param naming the field. Status codes and their meanings are shared across the API; see Errors.

The per-parameter contract this page summarizes is maintained in code, as PARAM_SUPPORT in minds/requests/chat_completions_request.py, and held to the OpenAI SDK's own parameter list by a test.