Models
You address models by alias: a short, stable name like sonnet or gpt. The alias resolves server-side to a concrete model on one of the upstream providers. When the underlying model is upgraded, the alias stays the same, so you don't have to chase version strings.
Jev uses the Decisions API, with native state and typed questions. Model catalog entries with kind: "decision" use /v1/decisions; the chat APIs described here apply to kind: "chat".
Listing models
GET /v1/models returns the live catalog with per-model metadata: label, whether your organization can call it right now, reasoning-effort levels, provider, and family. The row schema and the client-specific dialects are on the Models endpoint page; what each model supports is on the capability matrix.
The catalog
Current as of September 2026; GET /v1/models is authoritative and changes without notice. Prices are in Billing.
Three kinds of row, and the family field tells them apart:
- A moving alias (
family == id) tracks the newest release in its line. Its label names the line —GPT,Claude Sonnet— never the version it happens to serve, so the name stays put across repoints. What it resolves today can change; when it does, your requests move with it. - A fixed alias naming the model its moving alias serves today.
current_versionon the moving alias names it. Pick this one to hold a model through a repoint. - A fixed alias for an older release. Also carries
family, but is not named by its moving alias'scurrent_version.
Every model we serve is reachable by at least one alias that never moves.
| Alias | Label | Notes |
|---|---|---|
mindshub_air | MindsHub Air | Tracks the newest release. No fixed alias of its own. |
mindshub_blaze | MindsHub Blaze | Tracks the newest release. No fixed alias of its own. |
gpt-oss-fireworks | GPT-OSS 120B (Fireworks) | Fixed. mindshub_blaze's Fireworks-hosted fallback — a different model (GPT-OSS 120B) than mindshub_blaze currently serves. |
sonnet | Claude Sonnet | Tracks the newest release. Currently Claude Sonnet 5 — pin sonnet-5 to hold that model. |
sonnet-5 | Claude Sonnet 5 | Fixed. The same model sonnet serves today, and it stays on it when sonnet moves. |
opus | Claude Opus | Tracks the newest release. Currently Claude Opus 5 — pin opus-5 to hold that model. |
opus-5 | Claude Opus 5 | Fixed. The same model opus serves today, and it stays on it when opus moves. |
fable | Claude Fable | Tracks the newest release. Currently Claude Fable 5.1 — pin fable-5-1 to hold that model. |
fable-5-1 | Claude Fable 5.1 | Fixed. The same model fable serves today, and it stays on it when fable moves. |
fable-5 | Claude Fable 5 | Fixed. An older release of fable. |
haiku | Claude Haiku | Tracks the newest release. Currently Claude Haiku 4.5 — pin haiku-4-5 to hold that model. |
haiku-4-5 | Claude Haiku 4.5 | Fixed. The same model haiku serves today, and it stays on it when haiku moves. |
gpt | GPT | Tracks the newest release. Currently GPT-6 Astra — pin gpt-6-astra to hold that model. |
gpt-6-astra | GPT-6 Astra | Fixed. The same model gpt serves today, and it stays on it when gpt moves. |
gpt-5-6-sol | GPT 5.6 Sol | Fixed. An older release of gpt. |
gpt-terra | GPT Terra | Tracks the newest release. Currently GPT 5.6 Terra — pin gpt-terra-5-6 to hold that model. |
gpt-terra-5-6 | GPT 5.6 Terra | Fixed. The same model gpt-terra serves today, and it stays on it when gpt-terra moves. |
gpt-luna | GPT Luna | Tracks the newest release. Currently GPT 5.6 Luna — pin gpt-luna-5-6 to hold that model. |
gpt-luna-5-6 | GPT 5.6 Luna | Fixed. The same model gpt-luna serves today, and it stays on it when gpt-luna moves. |
gpt-codex | GPT Codex | Tracks the newest release. Currently GPT 5.3 Codex — pin gpt-codex-5-3 to hold that model. |
gpt-codex-5-3 | GPT 5.3 Codex | Fixed. The same model gpt-codex serves today, and it stays on it when gpt-codex moves. |
gpt-mini | GPT Mini | Tracks the newest release. Currently GPT 5.4 Mini — pin gpt-mini-5-4 to hold that model. |
gpt-mini-5-4 | GPT 5.4 Mini | Fixed. The same model gpt-mini serves today, and it stays on it when gpt-mini moves. |
gpt-nano | GPT Nano | Tracks the newest release. Currently GPT 5.4 Nano — pin gpt-nano-5-4 to hold that model. |
gpt-nano-5-4 | GPT 5.4 Nano | Fixed. The same model gpt-nano serves today, and it stays on it when gpt-nano moves. |
gemini | Gemini Pro | Tracks the newest release. Currently Gemini 3.1 Pro Preview — pin gemini-3-1-pro to hold that model. |
gemini-3-1-pro | Gemini 3.1 Pro Preview | Fixed. The same model gemini serves today, and it stays on it when gemini moves. |
gemini-flash | Gemini Flash | Tracks the newest release. Currently Gemini 3.8 Flash — pin gemini-flash-3-8 to hold that model. |
gemini-flash-3-8 | Gemini 3.8 Flash | Fixed. The same model gemini-flash serves today, and it stays on it when gemini-flash moves. |
gemini-flash-3-7 | Gemini 3.7 Flash | Fixed. An older release of gemini-flash. |
gemini-flash-3-6 | Gemini 3.6 Flash | Fixed. An older release of gemini-flash. |
gemini-flash-3-5 | Gemini 3.5 Flash | Fixed. An older release of gemini-flash. |
gemini-flash-3 | Gemini 3 Flash Preview | Fixed. An older release of gemini-flash. |
gemini-flash-lite | Gemini Flash-Lite | Tracks the newest release. Currently Gemini 3.1 Flash-Lite — pin gemini-flash-lite-3-1 to hold that model. |
gemini-flash-lite-3-1 | Gemini 3.1 Flash-Lite | Fixed. The same model gemini-flash-lite serves today, and it stays on it when gemini-flash-lite moves. |
kimi | Kimi | Tracks the newest release. Currently Kimi K3 — pin kimi-k3 to hold that model. |
kimi-k3 | Kimi K3 | Fixed. The same model kimi serves today, and it stays on it when kimi moves. |
deepseek-v4-flash | DeepSeek Flash | Tracks the newest release. Currently DeepSeek V4.1 Flash. Pin deepseek-v4-1-flash to hold that model. |
deepseek-v4-1-flash | DeepSeek V4.1 Flash | Fixed. The same model deepseek-v4-flash serves today, and it stays on it when deepseek-v4-flash moves. |
qwen | Qwen | Tracks the newest release. Currently Qwen3.8-2.4T-A95B — pin qwen-3-8-a95b to hold that model. |
qwen-3-8-a95b | Qwen3.8-2.4T-A95B | Fixed. The same model qwen serves today, and it stays on it when qwen moves. |
glm | GLM | Tracks the newest release. Currently GLM 5.3 — pin glm-5-3 to hold that model. |
glm-5-3 | GLM 5.3 | Fixed. The same model glm serves today, and it stays on it when glm moves. |
glm-5-3-flash | GLM Flash | Tracks the newest release. Currently GLM 5.3 Flash — pin glm-flash-5-3 to hold that model. |
glm-flash-5-3 | GLM 5.3 Flash | Fixed. The same model glm-5-3-flash serves today, and it stays on it when glm-5-3-flash moves. |
muse-spark | Muse Spark | Tracks the newest release. Currently Muse Spark 1.3 — pin muse-spark-1-3 to hold that model. |
muse-spark-1-3 | Muse Spark 1.3 | Fixed. The same model muse-spark serves today, and it stays on it when muse-spark moves. |
muse-spark-1-2 | Muse Spark 1.2 | Fixed. An older release of muse-spark. |
muse-spark-1-1 | Muse Spark 1.1 | Fixed. An older release of muse-spark. |
grok | Grok | Tracks the newest release. Currently Grok 4.7 — pin grok-4-7 to hold that model. |
grok-4-7 | Grok 4.7 | Fixed. The same model grok serves today, and it stays on it when grok moves. |
grok-4-6 | Grok 4.6 | Fixed. An older release of grok. |
grok-4-5 | Grok 4.5 | Fixed. An older release of grok. |
jev | Jev | Decision model. Currently jev-1.13.0; jev-latest is an accepted synonym. Uses /v1/decisions. |
jev-1.13.0 | Jev 1.13.0 | Fixed decision-model version. |
embed-small | Text Embedding | Tracks the newest release. Currently Text Embedding 3 (small) — pin embed-3-small to hold that model. |
embed-3-small | Text Embedding 3 (small) | Fixed. The same model embed-small serves today, and it stays on it when embed-small moves. |
Alias rules
- Send the bare alias.
"model": "sonnet". - Provider model IDs work too.
"model": "claude-sonnet-5"or"gpt-5.4-mini"resolves to the catalog entry serving that model, which is what makes a response'smodelsafe to send back. The Anthropic-compatible/v1/messagesendpoint also maps Claude family names onto aliases so Claude Code works unmodified; see Messages. latest:<alias>is deprecated but still accepted;latest:sonnetis exactlysonnet. New code should not use it.- Unknown names return
404with error codemodel_not_found, whether the name is a typo, a raw provider ID, or a model that exists but isn't in the catalog.
What the response's model field contains
Chat and embedding endpoints name the model that served the request. Ask for haiku and the response says claude-haiku-4-5-20251001. If the requested model was down and a fallback answered, model names the fallback. This holds on /v1/chat/completions, /v1/responses, /v1/messages and /v1/embeddings alike, on streaming frames as well as whole responses.
mindshub_air and mindshub_blaze are the exception. They return the string you sent (mindshub_air), fallback included. The model behind them is ours to move, so it is never named.
The value is safe to reuse: send it back as the model of your next request and it resolves. It resolves to the alias that currently serves that model, not to a pin: send back claude-haiku-4-5-20251001 and you are on haiku, which moves when haiku is repointed. To hold one model, send its pinned alias from the catalog (haiku-4-5).
The Decisions endpoint returns the actual served Jev version instead: requesting jev currently returns jev-1.13.0.
For chat and embeddings, three consequences are worth knowing:
- Log the alias you requested alongside the returned
model. The returned value changes when we repoint an alias or a fallback answers, so aggregate on your own requested alias if you want a stable key. - A fallback is billed at the rate of the alias you requested, even when
modelnames a different model. See Billing. - On
mindshub_airandmindshub_blaze,modeldoes not reveal a fallback. Read the failover signal instead.
Failover
When the requested model fails with a retryable error (a timeout, a 429, most 5xx), the request is retried on the same model and then handed to a fallback. Every response from chat completions, responses and messages says whether that happened, on success and on error, streamed or not, with one exception described below:
| Header | Value |
|---|---|
X-MindsHub-Failover | true when a fallback produced the response, false when the requested model did. |
X-MindsHub-Attempts | Upstream attempts, requested model and fallbacks together. 1 means the first try answered; 2 with false means one retry on the requested model. It counts the platform's own attempts, not retries inside a provider SDK. |
The same two values are in the body, for SDKs that only expose headers through a raw-response call: "mindshub": {"failover": false, "attempts": 1} at the top level of a whole response, and on the last event of a stream (the final chat chunk before [DONE], the closing Anthropic message_delta, or the response in response.completed, response.incomplete or response.failed). Error bodies don't carry it; read the headers there. Neither signal names the fallback model or its provider.
Both headers are present on every response that reached a model, except the stream described next, so a missing header on any other response means the request was refused before any model ran (a bad key, an exhausted allowance), not that no failover happened.
A stream whose 200 went out early carries neither header. A streamed request whose response takes more than 90 seconds to start gets its 200 before the response exists (see Streaming), so it carries none of the X-MindsHub-* headers, and a text/event-stream response without them is one of these. Its body still reports failover: the mindshub field rides its last event as usual, a response.failed on Responses included. A Chat Completions or Messages stream that fails after its early 200 ends in an error frame without mindshub, so that failure reports no failover signal at all.
Turning failover off. Send X-MindsHub-Allow-Failover: false to be served only by the model you asked for. Retries on that model still happen and are counted; if they all fail, you get its own error instead of a fallback's answer. Use it for evals, benchmarks, or a data-handling rule that allows one provider only. Any other value, or no header, keeps failover on.
Reasoning effort
Models with a non-null reasoning_efforts list accept an effort level; the levels vary by model, a level a model can't take is clamped onto its ladder rather than failing the request, and reasoning_efforts: null means the level isn't tunable, not that the model won't reason. Full rules and samples in Reasoning.
Model behavior differences
Wire behavior varies by provider: streaming chunk details, tool_choice handling, finish_reason mapping, and which sampling parameters are honored all differ. The per-model table is the capability matrix; the per-feature notes are in each guide.