Skip to main content

Rate limits

Two separate mechanisms can stop a request, and they return different errors:

  1. Throughput limits: how fast you can go. Hitting one returns 429 rate_limited; slowing down fixes it.
  2. Funding: whether your organization's included allowance or wallet can fund the request. Running out returns 429 included_allowance_exhausted or 402 wallet_empty, and a pause on free serving returns 429 free_air_daily_spend_fuse_exceeded; slowing down does not fix these.

Throughput limits​

Limits apply per organization (not per key, not per user), across all models combined:

LimitDefault
Requests per minute60
Tokens per minute1,000,000
Concurrent in-flight requests20

These are the standard defaults; they can be raised per organization; contact support@mindsdb.com if your workload needs more. There is currently no API that reports your organization's limits or remaining headroom; the only signal is the 429 itself.

Exceeding any of the three returns the same error:

HTTP 429
Retry-After: 12

{"error": {"message": "Rate limit exceeded for model 'sonnet'. Please slow down and retry.",
"type": "rate_limit_error", "param": null, "code": "rate_limited"}}

Honor Retry-After (seconds, always ≥ 1). The response does not identify which of the three limits you exceeded; if you see 429s at low request rates, suspect the token budget or the concurrency cap.

How the token budget is counted​

When a request is admitted, the per-minute token budget reserves the larger of your estimated prompt size and your max_tokens, then settles to the real count after the request finishes. Consequences:

  • A request with a huge max_tokens reserves that many tokens from the minute's budget up front, even if the model would have answered in 50. Oversized max_tokens values directly reduce how many requests you can run per minute.
  • A burst of concurrent large-max_tokens requests can exhaust the budget before any of them finishes.
  • A single request bigger than the whole per-minute budget isn't rejected outright; it waits for the budget to refill, surfacing as 429s until then.

Keep max_tokens realistic, but not too tight on models that reason internally; see the note in Chat completions.

No rate-limit headers on success​

Successful responses carry no X-RateLimit-* headers. Build client pacing on the 429s and Retry-After, not on header telemetry.

Running out of tokens instead​

Distinct from throughput: your organization's recurring included allowance can run out, the wallet can run dry, and free serving can be paused for everyone. Retrying does not help with any of them:

  • 429 included_allowance_exhausted: your own allowance is used up. Wait for reset_at in the body or X-MindsHub-Reset-At in the headers, or add credit. When neither is present, your organization's included allowance is zero (never granted, or set to zero), so waiting will not help and only credit will.
  • 429 free_air_daily_spend_fuse_exceeded: free serving is paused for everyone, even when your own allowance is untouched. The same reset_at and X-MindsHub-Reset-At say when free serving resumes. Wait until then, or add credit to continue immediately.
  • 402 wallet_empty: the request needs wallet credit. Add credit, then send it again.

Neither billing 429 sends Retry-After. Both send x-should-retry: false, so the OpenAI and Anthropic SDKs hand them to you on the first attempt instead of retrying. Billing shows how to read the remaining percentage; Errors has the exact shapes.

Practical guidance​

  • Cap client-side concurrency at or below 20, and queue beyond it.
  • Retry 429 rate_limited honoring Retry-After. Stop on included_allowance_exhausted and free_air_daily_spend_fuse_exceeded, and on any response that carries x-should-retry: false. A worked loop is in Errors.
  • Don't rely on the limiter as flow control: under degraded operation limits may not be enforced exactly, and nothing guarantees the 429 arrives precisely at the documented threshold.