Skip to main content

Errors

Every error the API returns, and what to do about each. Debugging a live failure? Skip to the error catalog; the body shapes below explain why the JSON you got may not match what your SDK expects.

For Jev, use the Decisions error and retry guidance. Its provider errors retain their native bodies, and an ambiguous timeout is not retried automatically.

Error shapes​

Every operation's error responses, with these shapes and the headers each carries, are also listed on its reference page.

The shape depends on where the error happens, not just what it is. There are four:

1. The standard error object: most errors on the OpenAI-compatible endpoints (/v1/chat/completions, /v1/responses, /v1/embeddings, /v1/models):

{
"error": {
"message": "The model 'not-a-model' does not exist or you do not have access to it.",
"type": "invalid_request_error",
"param": "model",
"code": "model_not_found"
}
}

A denial with a named cause adds deny_detail inside error, beside code. Today the only value is model_restricted, on a 403 permission_denied for a model an administrator in your organization restricted:

{
"error": {
"message": "An administrator in your organization has restricted the model 'opus'. Choose another model.",
"type": "permission_error",
"param": "model",
"code": "permission_denied",
"deny_detail": "model_restricted"
}
}

2. The Anthropic envelope: everything on /v1/messages and /v1/messages/count_tokens:

{"type": "error", "error": {"type": "not_found_error", "message": "..."}}

The same deny_detail appears inside error in this envelope too.

3. Plain {"detail": ...}: request-body validation failures (422 with a list of field errors, plus some 400s such as an invalid image data URL) and internal server errors (500) on most OpenAI-compatible endpoints. An OpenAI SDK will not find an error.message in these; read detail. (Exception: /v1/embeddings internal errors use {"error": {"type": "internal_error", ...}}, and its upstream provider errors are relayed with their original OpenAI-shaped bodies.)

4. Gate refusals: your credentials are checked in front of the API, before the endpoint is chosen, so a 401 carries the Anthropic envelope above on every endpoint, including the OpenAI-compatible ones:

{"type": "error", "error": {"type": "authentication_error", "message": "Invalid or missing credentials. See https://docs.mindshub.ai/inference/authentication"}}

A gate 403 arrives in the same shape with error.type of permission_error, and means the credential is real but is not allowed to use that resource. Dispatch on the status code; the message is written for a person and may change. See Authentication.

Credentials are checked twice: once at the edge, and once at the deeper authorization gate the request passes on its way to a model. Both refusals look the same to you, and both are terminal.

Dispatch on the HTTP status code, and treat the body as best-effort. The code field, where present, is stable and safe to switch on.

Error catalog​

StatuscodeWhat happenedWhat to do
401none or invalid_credentialsMissing, malformed, or revoked credentials: an API key or a browser session. The code is absent when the edge refused the credential and invalid_credentials when the authorization gate did; both mean the same thing to you.Fix the credentials. Don't retry.
403none or permission_deniedThe credential is real but is not allowed to use this resource, for example a removed organization member or a role without model execution permission.Fix the credential's access. Don't retry. If you believe the credential should have access, report it to support@mindsdb.com with the response headers.
403permission_denied with deny_detail: model_restrictedAn administrator in your organization restricted this model. The body's error.deny_detail and the X-MindsHub-Deny-Detail header both read model_restricted, and the message names the model. GET /v1/models lists it with enabled: false and disabled_reason: model_restricted.Choose another model. Don't retry, and don't add credit: credit does not lift an administrator's rule.
402wallet_emptyThe request needs wallet credit and your organization's balance is empty. This also applies to free-bucket overage when the account should recover through its wallet, and, on /v1/decisions, to an organization with neither credit nor a topped-up card free of payment errors, once its free daily Jev allowance or the shared free capacity is used up.Add credit in the console; the X-MindsHub-Recovery-Url header carries the console path. On a priced model, don't retry until funded. On /v1/decisions the call also works again without credit once the free allowance refills, gradually over 24 hours. See Billing.
400max_tokens_exceededmax_tokens above the hard cap of 131,072.Lower max_tokens.
400model_not_configuredThe model is in the catalog but isn't currently routable. Rare.Use another model, and report it to support@mindsdb.com.
404model_not_foundUnknown model name: a typo, a raw provider ID, or an alias not in the catalog.Use an alias from GET /v1/models.
422none (detail list)Request body failed validation: missing required field, wrong type, unknown message role.Fix the request.
429rate_limitedToo fast: requests per minute, tokens per minute, or concurrent requests.Wait and retry, honoring Retry-After (seconds, always ≥ 1). See Rate limits.
429included_allowance_exhaustedYour weighted allowance is used up and no wallet capacity is available. An organization whose included allowance is zero (never granted, or set to zero) gets this too, with no reset_at, because nothing refills.Wait for reset_at in the body or X-MindsHub-Reset-At in the headers, or add credit. When neither is present, only credit helps. No Retry-After is sent. Carries x-should-retry: false, so the OpenAI and Anthropic SDKs return it on the first attempt.
429free_air_daily_spend_fuse_exceededFree serving is paused for everyone until the daily budget resets. Your own allowance may be untouched.Wait for reset_at, which is the next UTC midnight, or add credit to continue immediately. No Retry-After is sent. Carries x-should-retry: false, so the OpenAI and Anthropic SDKs return it on the first attempt.
4xx (relayed)variesThe upstream provider rejected something the platform forwarded, for example a temperature the model doesn't accept. Body has "type": "api_error" and the provider's own message and status.Fix the request for that model, or switch models.
5xx (relayed)variesThe upstream provider failed. Body has "type": "api_error". The platform retries transient provider errors (and fails over where a fallback route exists) before you see this — except where the response carries x-should-retry: false, which marks a failure retrying cannot change.Retry with backoff, unless x-should-retry: false is set.
504assembly_timeoutA non-streaming request ran past its 110-second limit. The limit covers the whole request, failover attempts and web-search turns included, because the edge in front of the API closes a connection that sends nothing for 125 seconds. It usually follows a long reasoning run or a long answer. If MindsHub was running a search or MCP loop for the request (the External search loop and our client Remote MCP columns under Providers), the turns, searches and MCP calls that finished before the cut are metered. On Claude models with a large max_tokens, the prompt is metered, but not the output generated before the cut. Other timed-out calls record no tokens.Don't retry: it will time out the same way, and may bill again. Resend with "stream": true, which is not held to this limit (see Streaming), or lower max_tokens (Claude, Fireworks-hosted models) or reasoning_effort (Grok, GPT).
502noneCouldn't reach the upstream provider.Retry with backoff.
503policy_unavailableMindsHub couldn't verify your account's access, so the request was refused instead of run. Transient.Retry in a few seconds.
500none (detail)Unhandled internal error.Retry once; if it persists, report it to support@mindsdb.com.

The 502 and 504 rows above arrive as a 503. The network edge in front of the gateway replaces a 502 or 504 with its own error page, discarding the body and headers, so the gateway sends them as a 503 instead, with the original status in error.upstream_status (at the top level of the body when there is no error object). A timed-out request therefore arrives as a 503 with code: assembly_timeout, upstream_status: 504 and x-should-retry: false, and a relayed provider 502 or 504 arrives as a 503 carrying that status in upstream_status. A 502 or 504 you do receive came from infrastructure in front of the gateway, not from the gateway itself. The OpenAI and Anthropic SDKs raise the same exception for all three statuses and retry them the same way.

The three 429s mean different things. rate_limited means slow down and retry after seconds. included_allowance_exhausted means your own allowance is gone; wait until reset_at or top up. If it arrives with no reset_at, your organization's included allowance is zero, so waiting will not help and adding credit will. free_air_daily_spend_fuse_exceeded means free serving is paused for everyone until the next UTC day, so retrying sooner will not help and adding credit will. Only the rate-limit response sends Retry-After. The other two send x-should-retry: false, so an OpenAI or Anthropic SDK retries rate_limited on its own and hands you the other two on the first attempt.

The X-MindsHub-* headers​

Denials from access checks carry machine-readable headers that survive even if an intermediary rewrites the response body:

HeaderSent onValue
X-MindsHub-Reasoninvalid_credentials, permission_denied, model_not_found, wallet_empty, included_allowance_exhausted, free_air_daily_spend_fuse_exceeded, rate_limited, policy_unavailableThe denial reason. Usually matches the body code, except model_not_found, where the header reads unknown_model.
X-MindsHub-Deny-Detailpermission_denied for a model an administrator in your organization restrictedmodel_restricted. Same value as the body's error.deny_detail. Absent on every other permission_denied.
Retry-Afterrate_limited onlySeconds to wait, always ≥ 1.
X-MindsHub-Reset-Atincluded_allowance_exhausted, free_air_daily_spend_fuse_exceededSame ISO-8601 instant as the body-level reset_at: when your allowance refills, or when free serving resumes. Absent on included_allowance_exhausted when your organization's included allowance is zero.
X-MindsHub-Recovery-Urlwallet_emptyConsole path for adding credit. Not sent on model_restricted, which credit cannot fix.
x-should-retryincluded_allowance_exhausted, free_air_daily_spend_fuse_exceeded, assembly_timeoutfalse: the status is one an SDK would retry, but retrying fails the same way. The OpenAI and Anthropic SDKs check it before the status.

Two more X-MindsHub-* headers travel on successful responses rather than errors, reporting how a request was adapted to the target model. (X-MindsHub-Failover and X-MindsHub-Attempts, which say whether a fallback served the request, are on both, except on the stream described below; see Failover.)

HeaderSent onValue
X-MindsHub-Dropped-ParamsAny request carrying a parameter the model can't takeComma-separated parameter names, e.g. temperature,top_p
X-MindsHub-Clamped-ParamsAny request carrying a parameter outside the model's rangename=requested>applied, e.g. max_tokens=200000>128000

Neither appears when nothing was adapted, so their absence means the request went upstream as you wrote it, except on the stream described next. This is why a parameter a model dislikes usually produces a 200 rather than a 400; see Chat completions. A rewritten tool_choice is reported here too, as a clamp: tool_choice=required>auto (why the rewrite happens).

A stream whose 200 went out early carries none of the X-MindsHub-* headers. When a streamed response takes more than 90 seconds to start, its 200 goes out before the response exists (see Streaming), so no X-MindsHub-* header is sent, whatever was adapted, whether a fallback served, or whether a previous_response_id chain was cut short. On such a stream their absence says nothing, and nothing in the body reports the adaptations or a cut chain either. You can tell such a stream apart: it is a text/event-stream response without X-MindsHub-Failover, a header every other stream carries. Its failover signal still arrives in the body, as the mindshub field of its last event, a response.failed on Responses included. A Chat Completions or Messages stream that fails after its early 200 reports no failover signal at all (see Failover).

One more gap to code around: request-validation errors (max_tokens_exceeded, model_not_configured) carry none of these headers. Everywhere else a denial keeps its Retry-After and its X-MindsHub-* headers, /v1/messages included, so back off on the value we send rather than guessing. See Anthropic compatibility → Errors for that endpoint's envelope quirks.

Streaming failures​

A request with "stream": true can fail three ways:

  1. Before any output: you get a normal JSON error response (any of the above) instead of an SSE stream. Check the response's Content-Type before parsing it as SSE. The exception is a response that took more than 90 seconds to start, which gets its 200 early to keep the connection open. A failure after that arrives inside the stream: a data: frame with an error object on Chat Completions, response.created then response.failed on Responses, and an error event on Messages. Your SDK does not retry a failure that arrives this way, and the failure comes without its own HTTP status or headers. Retry it yourself where you would have retried that failure's status, and tell which failure it is from the error object on Chat Completions, response.error.code on Responses, or error.type on Messages.
  2. Mid-stream, visibly: the stream ends abnormally. There is no error event. On the OpenAI-compatible endpoints, a stream that ends without ever delivering a finish_reason chunk was a failed generation.
  3. Mid-stream, invisibly: on /v1/messages, a failed generation can arrive looking like a normal completion: the stream closes with an ordinary message_delta (stop_reason: "end_turn") and message_stop, just with truncated or empty content. There is currently no wire-level signal for this case. If your application must detect it, sanity-check the output (for example, empty content on a prompt that should produce text).

One model-family caveat for rule 2: streams on Gemini-served models (gemini, gemini-flash) always end with a finish_reason, even on failure, so on those, like on /v1/messages, absence of finish_reason can't be your only health check.

Partial output that arrived before a failure is real output and is billed.

A retry recipe​

Using plain HTTP (the requests library) so the status and headers are visible; SDK users get the same decisions from their SDK's status-code exceptions:

import random
import time
import requests

class OutOfTokens(Exception): pass
class RequestFailed(Exception): pass

RETRYABLE = {429, 500, 502, 503, 504}

def server_says_dont_retry(response):
# Some failures carry a retryable status but cannot be fixed by retrying.
# An `assembly_timeout` re-runs the same generation and bills again, and
# the two billing 429s meet the same limit again. Both the Anthropic and
# OpenAI SDKs check this header for the same reason.
return response.headers.get("x-should-retry") == "false"

def exhausted_allowance(response):
# Prefer the header; fall back to the body's error code
# (the header isn't sent on every endpoint or error).
if response.headers.get("X-MindsHub-Reason") == "included_allowance_exhausted":
return True
try:
return response.json().get("error", {}).get("code") == "included_allowance_exhausted"
except ValueError:
return False

def post_with_retries(url, headers, body, max_attempts=5, timeout=120):
for attempt in range(max_attempts):
try:
response = requests.post(url, headers=headers, json=body, timeout=timeout)
except requests.RequestException:
# Connection failure or timeout: the request may still have run
# server-side. See the note on duplicates below.
time.sleep(min(2 ** attempt, 30) + random.random())
continue

if response.ok:
return response
if exhausted_allowance(response):
raise OutOfTokens(response.headers.get("X-MindsHub-Reset-At"))
if response.status_code not in RETRYABLE or server_says_dont_retry(response):
raise RequestFailed(response.status_code, response.text)

retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else min(2 ** attempt, 30)
time.sleep(max(delay, 1.0) + random.random())

raise RequestFailed("gave up", max_attempts)

The decisions that matter: retry 429 rate_limited and 5xx with backoff and jitter; honor Retry-After when present; honor x-should-retry: false; stop when the allowance is exhausted until reset_at, or until you add credit when no reset_at is sent; and never blind-retry 401, 402, 403, 404, 422, or a caller-caused 400.

This recipe is written for the OpenAI-compatible endpoints. The Anthropic-compatible /v1/messages envelope has no code, but allowance exhaustion still carries reset_at at the top level and X-MindsHub-Reset-At in the headers. Use either value to stop retrying until the reset. When neither is sent, your organization's included allowance is zero, and only credit lifts the denial. Both billing 429s carry x-should-retry: false on this endpoint too, and a model an administrator restricted carries error.deny_detail and X-MindsHub-Deny-Detail.

Retrying after a timeout can generate twice. A request that timed out client-side may still complete server-side and meter, and there is no idempotency key to deduplicate it. For large generations, prefer a generous timeout over aggressive retry, and treat a timeout as "outcome unknown", not "failed".