Skip to main content

Core concepts

A handful of ideas explain most of how MindsHub behaves.

Aliases​

You address models by a short, stable name (sonnet, gpt, kimi), not by a provider version string. The alias resolves server-side to a concrete model at the provider.

{"model": "sonnet"}

Aliases exist so that a provider version bump doesn't become a deploy on your side. When Claude Sonnet 5 is superseded, sonnet follows it and your code doesn't change.

Two consequences:

  • Provider model IDs also resolve, so claude-sonnet-5 works as model. Messages additionally maps Claude family names onto aliases so Claude Code's picker works unmodified.
  • Responses name the model that served, except on mindshub_air and mindshub_blaze, which return their own alias. Decisions returns the served Jev version. Log both your requested alias and the returned model.

The live list is GET /v1/models; the catalog is in Models.

Passthrough​

MindsHub does not host models. Each request is translated into the target provider's format, sent to that provider, and translated back into the format you asked in. There is no MindsHub model sitting in the middle rewriting your prompts.

This is why capability tracks the underlying model: if a model can't see images, MindsHub can't change that. The catalog identifies the configured provider. Jev runs at TypeSafe and preserves its native question-and-answer format through the Decisions endpoint.

Three APIs, one engine​

Chat Completions, Responses, and Messages are three wire formats over the same inference pipeline. Any chat model is reachable from all three. Pick the one your existing code speaks.

The formats differ in shape, not in power. Where a genuine capability gap exists (explicit cache breakpoints, the token-counting endpoint), it's flagged in the capability matrix.

Decision models use their own /v1/decisions endpoint. They return categories, probabilities, and scores from shared state. Embedding models use /v1/embeddings.

Parameters are adapted, not rejected​

The following rules apply to the three chat APIs. Decisions preserves its native fields and uses provider validation.

Models disagree about which generation parameters they accept. MindsHub is a router, so instead of failing a request over a parameter the target model doesn't take:

  • A parameter the model supports is passed through at whatever value you sent.
  • A parameter it doesn't support is dropped, and named in X-MindsHub-Dropped-Params.
  • A value above the model's range (max_tokens over its ceiling, a reasoning_effort above its ladder) is clamped down, and named in X-MindsHub-Clamped-Params as name=requested>applied.
  • An unknown top-level field is ignored silently.

Neither header appears when nothing changed, or on a stream whose 200 went out early, which carries no X-MindsHub-* header whatever changed (see Errors). SDKs and coding agents treat a 400 as a hard failure, which is why unsupported parameters are dropped rather than bounced.

The boundaries:

  • A model can restrict the values of a parameter it supports, and that error passes through as the provider's 400: Kimi K3 takes temperature only at 1 and top_p only at 0.95.
  • A model can refuse two parameters it supports when they arrive together. Claude Haiku 4.5 (haiku, haiku-4-5) takes temperature or top_p, not both: when both are sent, top_p is dropped and named in X-MindsHub-Dropped-Params. The capability matrix notes each such pair.
  • Unknown nested content, such as an unrecognized block inside messages, can still reach the provider and fail there.
  • Structured output is the one exception to the drop rule — see below.

Structured output​

Ask for a JSON schema and you get JSON matching it, on all three APIs: response_format on Chat Completions, text.format on Responses, output_config.format on Messages. This is the one parameter that is refused rather than dropped. A model that cannot constrain its output returns 400 param_not_supported, because prose handed to a caller who is about to json.loads() it fails later and less legibly than an error here. Samples, the schema-less mode, and the per-model rules are in Structured output.

Reasoning effort​

Many current models reason before answering. Where it's adjustable, it's a parameter (reasoning_effort, reasoning.effort, output_config.effort) whose levels vary by model; GET /v1/models is authoritative. Where it isn't adjustable, it still happens: mindshub_air and kimi reason internally on every request. Either way, reasoning bills as output tokens, so give reasoning models a few hundred tokens of max_tokens headroom. Levels, clamping, and extended thinking are in Reasoning.

Funding: included allowance and wallet​

Two separate pots pay for requests.

Included allowance: a recurring weighted allowance usable on mindshub_air. Cached reads consume less than uncached input, cache writes, or output. Read included_percent_remaining and next_refresh_at for the current window.

The wallet: a prepaid, organization-level balance that funds paid model requests, embeddings, and web search. On mindshub_air a cache write draws the allowance like any other prompt token, and a web search draws it too, and each only bills past it.

Jev is priced at $0 during its launch promotion and never draws the included allowance. An organization with wallet credit, or with a card that has completed a top-up and has no payment error, calls it without a daily cap. Every other organization gets a free daily allowance while shared free capacity lasts, and past it gets 402 wallet_empty. Permissions and rate limits apply to every organization. See Billing.

For paid requests and Jev decisions, funding is checked before the model runs, so an unfunded request is refused up front and costs nothing. Running out is not the same as going too fast:

MeaningFix
429 rate_limitedToo fastBack off, honor Retry-After
429 included_allowance_exhaustedIncluded allowance exhaustedWait for reset_at or add credit. No reset_at means your organization's included allowance is zero, so add credit
402 wallet_emptyWallet empty, or Jev's free daily allowance used upTop up. On Jev, the allowance also refills by itself over 24 hours

Only the first is fixed by retrying. Details in Billing and Rate limits.

Conversations live client-side, except on Responses​

On Chat Completions and Messages every request carries its full history. Responses is the exception: turns are stored by default and chain by previous_response_id, so a client can send only its new input. Resending history is cheaper than it looks because prompt caching is automatic across the catalog.

Terms in one line each​

TermMeaning
AliasShort stable model name you put in model
PassthroughRequests are translated and proxied to the upstream provider
Included allowanceRecurring weighted allowance on mindshub_air
WalletPrepaid balance funding everything else
Reasoning effortHow hard a model thinks before answering; billed as output
Dropped / clampedParameters adapted to the target model, reported in headers
Cache writePrompt tokens stored for reuse; draws the mindshub_air allowance at weight 12 and bills past it