Skip to main content

Models

You address models by alias: a short, stable name like sonnet or gpt. The alias resolves server-side to a concrete model on one of the upstream providers. When the underlying model is upgraded, the alias stays the same, so you don't have to chase version strings.

Jev uses the Decisions API, with native state and typed questions. Model catalog entries with kind: "decision" use /v1/decisions; the chat APIs described here apply to kind: "chat".

Listing models​

GET /v1/models returns the live catalog with per-model metadata: label, whether your organization can call it right now, reasoning-effort levels, provider, and family. The row schema and the client-specific dialects are on the Models endpoint page; what each model supports is on the capability matrix.

The catalog​

Current as of September 2026; GET /v1/models is authoritative and changes without notice. Prices are in Billing.

Three kinds of row, and the family field tells them apart:

  • A moving alias (family == id) tracks the newest release in its line. Its label names the line — GPT, Claude Sonnet — never the version it happens to serve, so the name stays put across repoints. What it resolves today can change; when it does, your requests move with it.
  • A fixed alias naming the model its moving alias serves today. current_version on the moving alias names it. Pick this one to hold a model through a repoint.
  • A fixed alias for an older release. Also carries family, but is not named by its moving alias's current_version.

Every model we serve is reachable by at least one alias that never moves.

AliasLabelNotes
mindshub_airMindsHub AirTracks the newest release. No fixed alias of its own.
mindshub_blazeMindsHub BlazeTracks the newest release. No fixed alias of its own.
gpt-oss-fireworksGPT-OSS 120B (Fireworks)Fixed. mindshub_blaze's Fireworks-hosted fallback — a different model (GPT-OSS 120B) than mindshub_blaze currently serves.
sonnetClaude SonnetTracks the newest release. Currently Claude Sonnet 5 — pin sonnet-5 to hold that model.
sonnet-5Claude Sonnet 5Fixed. The same model sonnet serves today, and it stays on it when sonnet moves.
opusClaude OpusTracks the newest release. Currently Claude Opus 5 — pin opus-5 to hold that model.
opus-5Claude Opus 5Fixed. The same model opus serves today, and it stays on it when opus moves.
fableClaude FableTracks the newest release. Currently Claude Fable 5.1 — pin fable-5-1 to hold that model.
fable-5-1Claude Fable 5.1Fixed. The same model fable serves today, and it stays on it when fable moves.
fable-5Claude Fable 5Fixed. An older release of fable.
haikuClaude HaikuTracks the newest release. Currently Claude Haiku 4.5 — pin haiku-4-5 to hold that model.
haiku-4-5Claude Haiku 4.5Fixed. The same model haiku serves today, and it stays on it when haiku moves.
gptGPTTracks the newest release. Currently GPT-6 Astra — pin gpt-6-astra to hold that model.
gpt-6-astraGPT-6 AstraFixed. The same model gpt serves today, and it stays on it when gpt moves.
gpt-5-6-solGPT 5.6 SolFixed. An older release of gpt.
gpt-terraGPT TerraTracks the newest release. Currently GPT 5.6 Terra — pin gpt-terra-5-6 to hold that model.
gpt-terra-5-6GPT 5.6 TerraFixed. The same model gpt-terra serves today, and it stays on it when gpt-terra moves.
gpt-lunaGPT LunaTracks the newest release. Currently GPT 5.6 Luna — pin gpt-luna-5-6 to hold that model.
gpt-luna-5-6GPT 5.6 LunaFixed. The same model gpt-luna serves today, and it stays on it when gpt-luna moves.
gpt-codexGPT CodexTracks the newest release. Currently GPT 5.3 Codex — pin gpt-codex-5-3 to hold that model.
gpt-codex-5-3GPT 5.3 CodexFixed. The same model gpt-codex serves today, and it stays on it when gpt-codex moves.
gpt-miniGPT MiniTracks the newest release. Currently GPT 5.4 Mini — pin gpt-mini-5-4 to hold that model.
gpt-mini-5-4GPT 5.4 MiniFixed. The same model gpt-mini serves today, and it stays on it when gpt-mini moves.
gpt-nanoGPT NanoTracks the newest release. Currently GPT 5.4 Nano — pin gpt-nano-5-4 to hold that model.
gpt-nano-5-4GPT 5.4 NanoFixed. The same model gpt-nano serves today, and it stays on it when gpt-nano moves.
geminiGemini ProTracks the newest release. Currently Gemini 3.1 Pro Preview — pin gemini-3-1-pro to hold that model.
gemini-3-1-proGemini 3.1 Pro PreviewFixed. The same model gemini serves today, and it stays on it when gemini moves.
gemini-flashGemini FlashTracks the newest release. Currently Gemini 3.8 Flash — pin gemini-flash-3-8 to hold that model.
gemini-flash-3-8Gemini 3.8 FlashFixed. The same model gemini-flash serves today, and it stays on it when gemini-flash moves.
gemini-flash-3-7Gemini 3.7 FlashFixed. An older release of gemini-flash.
gemini-flash-3-6Gemini 3.6 FlashFixed. An older release of gemini-flash.
gemini-flash-3-5Gemini 3.5 FlashFixed. An older release of gemini-flash.
gemini-flash-3Gemini 3 Flash PreviewFixed. An older release of gemini-flash.
gemini-flash-liteGemini Flash-LiteTracks the newest release. Currently Gemini 3.1 Flash-Lite — pin gemini-flash-lite-3-1 to hold that model.
gemini-flash-lite-3-1Gemini 3.1 Flash-LiteFixed. The same model gemini-flash-lite serves today, and it stays on it when gemini-flash-lite moves.
kimiKimiTracks the newest release. Currently Kimi K3 — pin kimi-k3 to hold that model.
kimi-k3Kimi K3Fixed. The same model kimi serves today, and it stays on it when kimi moves.
deepseek-v4-flashDeepSeek FlashTracks the newest release. Currently DeepSeek V4.1 Flash. Pin deepseek-v4-1-flash to hold that model.
deepseek-v4-1-flashDeepSeek V4.1 FlashFixed. The same model deepseek-v4-flash serves today, and it stays on it when deepseek-v4-flash moves.
qwenQwenTracks the newest release. Currently Qwen3.8-2.4T-A95B — pin qwen-3-8-a95b to hold that model.
qwen-3-8-a95bQwen3.8-2.4T-A95BFixed. The same model qwen serves today, and it stays on it when qwen moves.
glmGLMTracks the newest release. Currently GLM 5.3 — pin glm-5-3 to hold that model.
glm-5-3GLM 5.3Fixed. The same model glm serves today, and it stays on it when glm moves.
glm-5-3-flashGLM FlashTracks the newest release. Currently GLM 5.3 Flash — pin glm-flash-5-3 to hold that model.
glm-flash-5-3GLM 5.3 FlashFixed. The same model glm-5-3-flash serves today, and it stays on it when glm-5-3-flash moves.
muse-sparkMuse SparkTracks the newest release. Currently Muse Spark 1.3 — pin muse-spark-1-3 to hold that model.
muse-spark-1-3Muse Spark 1.3Fixed. The same model muse-spark serves today, and it stays on it when muse-spark moves.
muse-spark-1-2Muse Spark 1.2Fixed. An older release of muse-spark.
muse-spark-1-1Muse Spark 1.1Fixed. An older release of muse-spark.
grokGrokTracks the newest release. Currently Grok 4.7 — pin grok-4-7 to hold that model.
grok-4-7Grok 4.7Fixed. The same model grok serves today, and it stays on it when grok moves.
grok-4-6Grok 4.6Fixed. An older release of grok.
grok-4-5Grok 4.5Fixed. An older release of grok.
jevJevDecision model. Currently jev-1.13.0; jev-latest is an accepted synonym. Uses /v1/decisions.
jev-1.13.0Jev 1.13.0Fixed decision-model version.
embed-smallText EmbeddingTracks the newest release. Currently Text Embedding 3 (small) — pin embed-3-small to hold that model.
embed-3-smallText Embedding 3 (small)Fixed. The same model embed-small serves today, and it stays on it when embed-small moves.

Alias rules​

  • Send the bare alias. "model": "sonnet".
  • Provider model IDs work too. "model": "claude-sonnet-5" or "gpt-5.4-mini" resolves to the catalog entry serving that model, which is what makes a response's model safe to send back. The Anthropic-compatible /v1/messages endpoint also maps Claude family names onto aliases so Claude Code works unmodified; see Messages.
  • latest:<alias> is deprecated but still accepted; latest:sonnet is exactly sonnet. New code should not use it.
  • Unknown names return 404 with error code model_not_found, whether the name is a typo, a raw provider ID, or a model that exists but isn't in the catalog.

What the response's model field contains​

Chat and embedding endpoints name the model that served the request. Ask for haiku and the response says claude-haiku-4-5-20251001. If the requested model was down and a fallback answered, model names the fallback. This holds on /v1/chat/completions, /v1/responses, /v1/messages and /v1/embeddings alike, on streaming frames as well as whole responses.

mindshub_air and mindshub_blaze are the exception. They return the string you sent (mindshub_air), fallback included. The model behind them is ours to move, so it is never named.

The value is safe to reuse: send it back as the model of your next request and it resolves. It resolves to the alias that currently serves that model, not to a pin: send back claude-haiku-4-5-20251001 and you are on haiku, which moves when haiku is repointed. To hold one model, send its pinned alias from the catalog (haiku-4-5).

The Decisions endpoint returns the actual served Jev version instead: requesting jev currently returns jev-1.13.0.

For chat and embeddings, three consequences are worth knowing:

  • Log the alias you requested alongside the returned model. The returned value changes when we repoint an alias or a fallback answers, so aggregate on your own requested alias if you want a stable key.
  • A fallback is billed at the rate of the alias you requested, even when model names a different model. See Billing.
  • On mindshub_air and mindshub_blaze, model does not reveal a fallback. Read the failover signal instead.

Failover​

When the requested model fails with a retryable error (a timeout, a 429, most 5xx), the request is retried on the same model and then handed to a fallback. Every response from chat completions, responses and messages says whether that happened, on success and on error, streamed or not, with one exception described below:

HeaderValue
X-MindsHub-Failovertrue when a fallback produced the response, false when the requested model did.
X-MindsHub-AttemptsUpstream attempts, requested model and fallbacks together. 1 means the first try answered; 2 with false means one retry on the requested model. It counts the platform's own attempts, not retries inside a provider SDK.

The same two values are in the body, for SDKs that only expose headers through a raw-response call: "mindshub": {"failover": false, "attempts": 1} at the top level of a whole response, and on the last event of a stream (the final chat chunk before [DONE], the closing Anthropic message_delta, or the response in response.completed, response.incomplete or response.failed). Error bodies don't carry it; read the headers there. Neither signal names the fallback model or its provider.

Both headers are present on every response that reached a model, except the stream described next, so a missing header on any other response means the request was refused before any model ran (a bad key, an exhausted allowance), not that no failover happened.

A stream whose 200 went out early carries neither header. A streamed request whose response takes more than 90 seconds to start gets its 200 before the response exists (see Streaming), so it carries none of the X-MindsHub-* headers, and a text/event-stream response without them is one of these. Its body still reports failover: the mindshub field rides its last event as usual, a response.failed on Responses included. A Chat Completions or Messages stream that fails after its early 200 ends in an error frame without mindshub, so that failure reports no failover signal at all.

Turning failover off. Send X-MindsHub-Allow-Failover: false to be served only by the model you asked for. Retries on that model still happen and are counted; if they all fail, you get its own error instead of a fallback's answer. Use it for evals, benchmarks, or a data-handling rule that allows one provider only. Any other value, or no header, keeps failover on.

Reasoning effort​

Models with a non-null reasoning_efforts list accept an effort level; the levels vary by model, a level a model can't take is clamped onto its ladder rather than failing the request, and reasoning_efforts: null means the level isn't tunable, not that the model won't reason. Full rules and samples in Reasoning.

Model behavior differences​

Wire behavior varies by provider: streaming chunk details, tool_choice handling, finish_reason mapping, and which sampling parameters are honored all differ. The per-model table is the capability matrix; the per-feature notes are in each guide.