Core concepts
A handful of ideas explain most of how MindsHub behaves.
Aliases
You address models by a short, stable name (sonnet, gpt, kimi), not by a provider version string. The alias resolves server-side to a concrete model at the provider.
{"model": "sonnet"}
Aliases exist so that a provider version bump doesn't become a deploy on your side. When Claude Sonnet 5 is superseded, sonnet follows it and your code doesn't change.
Two consequences:
- Provider model IDs also resolve, so
claude-sonnet-5works asmodel. Messages additionally maps Claude family names onto aliases so Claude Code's picker works unmodified. - Responses name the model that served, except on
mindshub_airandmindshub_blaze, which return their own alias. Decisions returns the served Jev version. Log both your requested alias and the returnedmodel.
The live list is GET /v1/models; the catalog is in Models.
Passthrough
MindsHub does not host models. Each request is translated into the target provider's format, sent to that provider, and translated back into the format you asked in. There is no MindsHub model sitting in the middle rewriting your prompts.
This is why capability tracks the underlying model: if a model can't see images, MindsHub can't change that. The catalog identifies the configured provider. Jev runs at TypeSafe and preserves its native question-and-answer format through the Decisions endpoint.
Three APIs, one engine
Chat Completions, Responses, and Messages are three wire formats over the same inference pipeline. Any chat model is reachable from all three. Pick the one your existing code speaks.
The formats differ in shape, not in power. Where a genuine capability gap exists (explicit cache breakpoints, the token-counting endpoint), it's flagged in the capability matrix.
Decision models use their own /v1/decisions endpoint. They return categories, probabilities, and scores from shared state. Embedding models use /v1/embeddings.
Parameters are adapted, not rejected
The following rules apply to the three chat APIs. Decisions preserves its native fields and uses provider validation.
Models disagree about which generation parameters they accept. MindsHub is a router, so instead of failing a request over a parameter the target model doesn't take:
- A parameter the model supports is passed through at whatever value you sent.
- A parameter it doesn't support is dropped, and named in
X-MindsHub-Dropped-Params. - A value above the model's range (
max_tokensover its ceiling, areasoning_effortabove its ladder) is clamped down, and named inX-MindsHub-Clamped-Paramsasname=requested>applied. - An unknown top-level field is ignored silently.
Neither header appears when nothing changed, or on a stream whose 200 went out early, which carries no X-MindsHub-* header whatever changed (see Errors). SDKs and coding agents treat a 400 as a hard failure, which is why unsupported parameters are dropped rather than bounced.
The boundaries:
- A model can restrict the values of a parameter it supports, and that error passes through as the provider's 400: Kimi K3 takes
temperatureonly at1andtop_ponly at0.95. - A model can refuse two parameters it supports when they arrive together. Claude Haiku 4.5 (
haiku,haiku-4-5) takestemperatureortop_p, not both: when both are sent,top_pis dropped and named inX-MindsHub-Dropped-Params. The capability matrix notes each such pair. - Unknown nested content, such as an unrecognized block inside
messages, can still reach the provider and fail there. - Structured output is the one exception to the drop rule — see below.
Structured output
Ask for a JSON schema and you get JSON matching it, on all three APIs: response_format on Chat Completions, text.format on Responses, output_config.format on Messages. This is the one parameter that is refused rather than dropped. A model that cannot constrain its output returns 400 param_not_supported, because prose handed to a caller who is about to json.loads() it fails later and less legibly than an error here. Samples, the schema-less mode, and the per-model rules are in Structured output.
Reasoning effort
Many current models reason before answering. Where it's adjustable, it's a parameter (reasoning_effort, reasoning.effort, output_config.effort) whose levels vary by model; GET /v1/models is authoritative. Where it isn't adjustable, it still happens: mindshub_air and kimi reason internally on every request. Either way, reasoning bills as output tokens, so give reasoning models a few hundred tokens of max_tokens headroom. Levels, clamping, and extended thinking are in Reasoning.
Funding: included allowance and wallet
Two separate pots pay for requests.
Included allowance: a recurring weighted allowance usable on mindshub_air. Cached reads consume less than uncached input, cache writes, or output. Read included_percent_remaining and next_refresh_at for the current window.
The wallet: a prepaid, organization-level balance that funds paid model requests, embeddings, and web search. On mindshub_air a cache write draws the allowance like any other prompt token, and a web search draws it too, and each only bills past it.
Jev is priced at $0 during its launch promotion and never draws the included allowance. An organization with wallet credit, or with a card that has completed a top-up and has no payment error, calls it without a daily cap. Every other organization gets a free daily allowance while shared free capacity lasts, and past it gets 402 wallet_empty. Permissions and rate limits apply to every organization. See Billing.
For paid requests and Jev decisions, funding is checked before the model runs, so an unfunded request is refused up front and costs nothing. Running out is not the same as going too fast:
| Meaning | Fix | |
|---|---|---|
429 rate_limited | Too fast | Back off, honor Retry-After |
429 included_allowance_exhausted | Included allowance exhausted | Wait for reset_at or add credit. No reset_at means your organization's included allowance is zero, so add credit |
402 wallet_empty | Wallet empty, or Jev's free daily allowance used up | Top up. On Jev, the allowance also refills by itself over 24 hours |
Only the first is fixed by retrying. Details in Billing and Rate limits.
Conversations live client-side, except on Responses
On Chat Completions and Messages every request carries its full history. Responses is the exception: turns are stored by default and chain by previous_response_id, so a client can send only its new input. Resending history is cheaper than it looks because prompt caching is automatic across the catalog.
Terms in one line each
| Term | Meaning |
|---|---|
| Alias | Short stable model name you put in model |
| Passthrough | Requests are translated and proxied to the upstream provider |
| Included allowance | Recurring weighted allowance on mindshub_air |
| Wallet | Prepaid balance funding everything else |
| Reasoning effort | How hard a model thinks before answering; billed as output |
| Dropped / clamped | Parameters adapted to the target model, reported in headers |
| Cache write | Prompt tokens stored for reuse; draws the mindshub_air allowance at weight 12 and bills past it |