Rate limits
Two separate mechanisms can stop a request, and they return different errors:
- Throughput limits: how fast you can go. Hitting one returns
429 rate_limited; slowing down fixes it. - Funding: whether your organization's included allowance or wallet can fund the request. Running out returns
429 included_allowance_exhaustedor402 wallet_empty, and a pause on free serving returns429 free_air_daily_spend_fuse_exceeded; slowing down does not fix these.
Throughput limits
Limits apply per organization (not per key, not per user), across all models combined:
| Limit | Default |
|---|---|
| Requests per minute | 60 |
| Tokens per minute | 1,000,000 |
| Concurrent in-flight requests | 20 |
These are the standard defaults; they can be raised per organization; contact support@mindsdb.com if your workload needs more. There is currently no API that reports your organization's limits or remaining headroom; the only signal is the 429 itself.
Exceeding any of the three returns the same error:
HTTP 429
Retry-After: 12
{"error": {"message": "Rate limit exceeded for model 'sonnet'. Please slow down and retry.",
"type": "rate_limit_error", "param": null, "code": "rate_limited"}}
Honor Retry-After (seconds, always ≥ 1). The response does not identify which of the three limits you exceeded; if you see 429s at low request rates, suspect the token budget or the concurrency cap.
How the token budget is counted
When a request is admitted, the per-minute token budget reserves the larger of your estimated prompt size and your max_tokens, then settles to the real count after the request finishes. Consequences:
- A request with a huge
max_tokensreserves that many tokens from the minute's budget up front, even if the model would have answered in 50. Oversizedmax_tokensvalues directly reduce how many requests you can run per minute. - A burst of concurrent large-
max_tokensrequests can exhaust the budget before any of them finishes. - A single request bigger than the whole per-minute budget isn't rejected outright; it waits for the budget to refill, surfacing as 429s until then.
Keep max_tokens realistic, but not too tight on models that reason internally; see the note in Chat completions.
No rate-limit headers on success
Successful responses carry no X-RateLimit-* headers. Build client pacing on the 429s and Retry-After, not on header telemetry.
Running out of tokens instead
Distinct from throughput: your organization's recurring included allowance can run out, the wallet can run dry, and free serving can be paused for everyone. Retrying does not help with any of them:
429 included_allowance_exhausted: your own allowance is used up. Wait forreset_atin the body orX-MindsHub-Reset-Atin the headers, or add credit. When neither is present, your organization's included allowance is zero (never granted, or set to zero), so waiting will not help and only credit will.429 free_air_daily_spend_fuse_exceeded: free serving is paused for everyone, even when your own allowance is untouched. The samereset_atandX-MindsHub-Reset-Atsay when free serving resumes. Wait until then, or add credit to continue immediately.402 wallet_empty: the request needs wallet credit. Add credit, then send it again.
Neither billing 429 sends Retry-After. Both send x-should-retry: false, so the OpenAI and Anthropic SDKs hand them to you on the first attempt instead of retrying. Billing shows how to read the remaining percentage; Errors has the exact shapes.
Practical guidance
- Cap client-side concurrency at or below 20, and queue beyond it.
- Retry
429 rate_limitedhonoringRetry-After. Stop onincluded_allowance_exhaustedandfree_air_daily_spend_fuse_exceeded, and on any response that carriesx-should-retry: false. A worked loop is in Errors. - Don't rely on the limiter as flow control: under degraded operation limits may not be enforced exactly, and nothing guarantees the 429 arrives precisely at the documented threshold.