Billing
MindsHub Inference is pay-as-you-go, with no subscription. Two things fund your requests:
Included allowance. An organization with a free grant receives a recurring allowance for mindshub_air; one without a grant has an allowance of zero. It charges four token lanes at different weights, so efficient prompt caching preserves more useful work than uncached input or generated output. The window has a fixed duration and next_refresh_at gives its exact end. Programmatically, use included_percent_remaining; the old included_tokens object remains temporarily for compatibility and carries no token count despite its name.
The wallet. A prepaid, organization-level credit balance that funds paid models, embeddings, and web-search charges. Credits never expire. Jev is priced at $0 during its launch promotion; organizations without credit call it within a free daily allowance, described under Jev decisions.
A request passes the funding check when your wallet has available balance, when it uses mindshub_air with allowance remaining, or when it calls Jev and your organization has a card with a completed top-up and no payment error, or still has free Jev allowance left (Jev decisions). Authentication, access restrictions and rate limits still apply. Otherwise it is refused up front with 402 wallet_empty or 429 included_allowance_exhausted. Allowance exhaustion returns the same ISO-8601 reset_at in the JSON body and X-MindsHub-Reset-At header. An organization whose included allowance is zero (never granted, or set to zero) gets the same 429 with neither, because nothing refills; adding credit is the only way forward. Both billing 429s carry x-should-retry: false, so the OpenAI and Anthropic SDKs do not retry them. A refused request consumes nothing.
How the wallet works
- Top up from the console. Minimum top-up is $10.
- Auto-recharge (optional): set a threshold and a target: when the balance crosses the threshold, your card is charged back up to the target, typically within a few minutes. A monthly recharge cap (default $100) limits automatic charges; hitting it pauses recharging until the calendar month rolls over.
- Your card is only charged for top-ups and recharges. Usage consumes prepaid credits.
- A declined auto-recharge stops recharging until the billing owner fixes the payment method; it never retries a declined card on its own. Requests keep working until the balance runs out.
- Wallet visibility is role-based: the billing owner sees balances and dollar figures; other members see their own usage in tokens only.
- An empty wallet never revokes API keys (Authentication); requests resume automatically once credit lands.
What gets metered
Four token counters, per request:
| Counter | What it is | Billed at |
|---|---|---|
| Input tokens | Prompt tokens not served from cache | The model's input rate |
| Output tokens | Everything the model generates, including internal reasoning | The model's output rate |
| Cached input tokens | Prompt tokens read from the provider's prompt cache | The model's cached-input rate, roughly a tenth of input on most of the catalog but not uniformly so. Read it off the price list rather than deriving it |
| Cache-write tokens | Prompt tokens written into the cache | The model's cache-write rate: the input rate on most models, 1.25× input on the Claude family, mindshub_air, and the GPT 5.6 line. Not billed at all on the Gemini models or on mindshub_blaze, which shows n/a in the price list |
On chat models, prompt caching is automatic and you never have to ask for it. On most of the catalog MindsHub marks the provider's cache breakpoints on the stable prefix of your request; on the rest the provider caches implicitly and there is nothing for us to mark. A request that already carries its own cache_control breakpoints, as Claude Code sends on Messages, is passed through exactly as you sent it. You see the effect as cached_tokens in responses and a lower effective input price. One rough edge: on Claude targets a system prompt sent as an array of content parts rather than a plain string gets no breakpoint of its own. The rolling breakpoint on the previous message picks it up from the second turn onward, so the cost is one extra full-price turn rather than every turn.
How this interacts with the included allowance:
- On
mindshub_air, cached reads draw 1 weighted unit per token, uncached input draws 10, cache writes draw 12, and output draws 60. - A web search on
mindshub_airdraws 65,000 weighted units and bills to the wallet only once the allowance is spent. On its own that is room for about 100 searches per window; the search results the model reads also draw the allowance as input tokens, so a real session fits fewer. - The weights are allowance policy, not prices. Wallet billing still uses the raw token counters and the current price list.
- Page fetches, embeddings, and web searches on every other model always bill to the wallet, even when the request's model tokens were covered by the allowance.
The full matrix is the summary table at the bottom.
Price list
These are MindsHub's prices; they track the upstream providers' list prices except for Jev, which is priced at $0 during its launch promotion. Per million tokens, with rates from September 2026 (the console is the authoritative, current source):
A fixed alias naming the same model as its moving alias — gpt-6-astra against gpt, sonnet-5 against sonnet — is priced identically to it and is left out of this table. Pinning a model costs nothing; it only stops the alias moving. Aliases for older releases price separately and are listed.
| Model | Input | Output | Cached input | Cache write |
|---|---|---|---|---|
mindshub_air | $0.20 | $1.20 | $0.02 | $0.25 |
mindshub_blaze | $0.99 | $1.49 | $0.99 | n/a |
gpt-oss-fireworks | $0.15 | $0.60 | $0.02 | $0.15 |
haiku | $1.00 | $5.00 | $0.10 | $1.25 |
sonnet | $2.00 | $10.00 | $0.20 | $2.50 |
opus | $5.00 | $25.00 | $0.50 | $6.25 |
fable | $10.00 | $50.00 | $0.25 | $12.50 |
fable-5 | $10.00 | $50.00 | $1.00 | $12.50 |
gpt | $10.00 | $50.00 | $1.00 | $12.50 |
gpt-5-6-sol | $5.00 | $30.00 | $0.50 | $6.25 |
gpt-terra | $2.00 | $12.00 | $0.20 | $2.50 |
gpt-luna | $0.20 | $1.20 | $0.02 | $0.25 |
gpt-codex | $1.75 | $14.00 | $0.18 | $1.75 |
gpt-mini | $0.75 | $4.50 | $0.08 | $0.75 |
gpt-nano | $0.20 | $1.25 | $0.02 | $0.20 |
gemini | $2.00 | $12.00 | $0.20 | n/a |
gemini-flash | $0.75 | $3.75 | $0.08 | n/a |
gemini-flash-3-7 | $0.75 | $3.75 | $0.08 | n/a |
gemini-flash-3-6 | $0.75 | $3.75 | $0.08 | n/a |
gemini-flash-3-5 | $1.50 | $9.00 | $0.15 | n/a |
gemini-flash-3 | $0.50 | $3.00 | $0.05 | n/a |
gemini-flash-lite | $0.25 | $1.50 | $0.03 | n/a |
kimi | $3.00 | $15.00 | $0.30 | $3.00 |
deepseek-v4-flash | $0.22 | $0.66 | $0.01 | $0.22 |
qwen | $2.00 | $6.00 | $0.25 | $2.00 |
glm | $1.40 | $4.40 | $0.26 | $1.40 |
glm-5-3-flash | $0.15 | $0.50 | $0.03 | $0.15 |
muse-spark | $1.25 | $4.25 | $0.15 | $1.25 |
muse-spark-1-2 | $1.25 | $4.25 | $0.15 | $1.25 |
muse-spark-1-1 | $1.25 | $4.25 | $0.15 | $1.25 |
grok | $2.00 | $6.00 | $0.50 | $2.00 |
grok-4-6 | $2.00 | $6.00 | $0.50 | $2.00 |
grok-4-5 | $2.00 | $6.00 | $0.30 | $2.00 |
embed-small | $0.02 | n/a | n/a | n/a |
Jev decisions: jev, jev-1.13.0 | $0 (promotion) | $0 | n/a | n/a |
Aliases are described in Models. Jev also accepts jev-latest. Who can call Jev without a daily cap is under Jev decisions.
deepseek-v4-1-flash is the fixed alias for the model deepseek-v4-flash serves today, so it prices identically to it and is not listed separately above.
Jev decisions
Jev input and output are priced at $0 during its launch promotion, and Jev requests never draw the included allowance. How much you can call it depends on how your organization is funded:
- With wallet credit, or with a card that has completed at least one top-up and has no unresolved payment error, your organization calls Jev without a daily cap.
- Otherwise, your organization gets a free daily allowance of 100,000 tokens and 100 requests, while shared free capacity lasts.
The allowance refills continuously rather than resetting once a day: it adds that amount back over every 24 hours and never holds more than one day's worth, so a new organization, which starts with a full allowance, can use roughly twice that on its first day. A call that starts while the allowance has requests and tokens left runs to completion even if it uses more tokens than remain, and the overrun is taken from the refill that follows. Every call that passes the funding check counts as one request, including one the provider refuses or times out.
A card on its own does not lift the cap; a completed top-up does. Past the allowance, or when shared free capacity runs out first, POST /v1/decisions returns 402 wallet_empty before the model runs. Without credit, the call works again once the allowance refills. The X-MindsHub-Recovery-Url header carries the console path for adding credit, and credit lifts the cap within about a minute of landing. Organization access restrictions and rate limits apply to every organization. Usage is measured and returned on every call, and decision requests have no cache charges.
Long-context pricing
Some providers charge more for large prompts, and MindsHub passes those thresholds through. When the total prompt (uncached, cached, and cache-write tokens combined) exceeds the model's threshold, the higher rates apply to the entire request, on every dimension:
| Model | Threshold (prompt tokens) | Input / Output | Cached input | Cache write |
|---|---|---|---|---|
mindshub_air | above 272,000 | $0.40 / $1.80 | $0.04 | $0.50 |
gpt | above 272,000 | $20.00 / $75.00 | $2.00 | $25.00 |
gpt-5-6-sol | above 272,000 | $10.00 / $45.00 | $1.00 | $12.50 |
gpt-terra | above 272,000 | $4.00 / $18.00 | $0.40 | $5.00 |
gpt-luna | above 272,000 | $0.40 / $1.80 | $0.04 | $0.50 |
gemini | above 200,000 | $4.00 / $18.00 | $0.40 | n/a |
grok | at or above 200,000 | $4.00 / $12.00 | $1.00 | $4.00 |
grok-4-6 | at or above 200,000 | $4.00 / $12.00 | $1.00 | $4.00 |
grok-4-5 | at or above 200,000 | $4.00 / $12.00 | $0.60 | $4.00 |
Note the cached rates double too: a long, heavily-cached conversation is exactly the workload that crosses these thresholds. All other models price flat across their full context window.
Web search and fetch
Per 1,000 uses, on models where web tools are available. mindshub_blaze, gpt-oss-fireworks and embed-small are not among them: web tools are dropped on those aliases, so no search or fetch charge can arise on them no matter what your tools array says. On mindshub_air a search draws the included allowance first and bills at the rate below only past it.
| Models | Search | Fetch |
|---|---|---|
mindshub_air, sonnet, opus, fable, fable-5, haiku, gpt, gpt-5-6-sol, gpt-terra, gpt-luna, gpt-codex, gpt-mini, gpt-nano | $10.00 | n/a |
kimi, deepseek-v4-flash, qwen, glm, glm-5-3-flash, muse-spark, muse-spark-1-2, muse-spark-1-1 | $7.00 | $1.00 |
gemini, gemini-flash, gemini-flash-3-7, gemini-flash-3-6, gemini-flash-3-5, gemini-flash-3, gemini-flash-lite | $14.00 | n/a |
grok, grok-4-6, grok-4-5 | $5.00 | n/a |
Where no fetch price is listed, fetched pages bill as ordinary input tokens rather than per fetch.
Remote MCP
No per-call charge. The MCP server is yours and neither upstream charges to call it; what you pay is the tokens the server's tool schemas and results add to the turn, which usage reports as ordinary input and output tokens. Calls are counted (mcp_call) on your usage summary so the count is visible; the count carries no price today. See Remote MCP servers.
Checking usage and balance from code
Two account endpoints accept your API key. They live on auth.mindshub.ai (not api.mindshub.ai), and the trailing slash is part of the path.
GET https://auth.mindshub.ai/v1/entitlements/me/: what you can use right now:
{
"wallet": {"balance_usd": "90.45", "can_consume": true, "auto_recharge": {"enabled": true, "...": "..."}},
"is_billing_owner": true,
"models": [
{"id": "mindshub_air", "label": "MindsHub Air", "free_bucket": true, "embedding": false,
"enabled": true, "reasoning_efforts": null, "default_reasoning_effort": "medium"}
],
"included_percent_remaining": 42.5,
"next_refresh_at": "2026-08-25T18:20:26Z"
}
next_refresh_at is when the allowance refills. free_bucket marks the model it covers. included_percent_remaining is canonical and keeps fractional values, so 0.4% does not display as exhausted. The response also carries a deprecated included_tokens object for older clients, omitted from the example above and scheduled for removal. On a live allowance it restates the same proportion out of a limit of 100 and says so in a unit field; an uncapped allowance reports every counter as null, and an organization with no grant reports a limit of 0. Read included_percent_remaining instead. The absolute size of the allowance is not published.
GET https://auth.mindshub.ai/v1/usage/summary/?range=period&group_by=model: what you've consumed. Response (excerpt; some fields omitted for brevity):
{
"range": {"requested": "period", "start": "2026-07-01T00:00:00Z", "end": "2026-08-01T00:00:00Z"},
"results": [
{"dimensions": {"model_alias": "opus"},
"usage": {"input_tokens": 218809, "output_tokens": 226135,
"billable_input_tokens": 218809, "billable_output_tokens": 226135,
"search_count": 87, "fetch_count": 16},
"cost_usd": "19.54"}
],
"next_cursor": "eyJvIjogM30=",
"totals": {"usage": {"...": "..."}, "cost_usd": "19.54"}
}
Reading it correctly:
- This endpoint reports paid consumption only.
range=periodis the monthly paid billing period. The allowance runs on a separate fixed-duration window and is not reported here at all. Read it fromGET /v1/entitlements/me/, whereincluded_percent_remaininggives the proportion left andnext_refresh_atgives the instant the window refills. cost_usdis a decimal string, populated only for the billing owner (nullotherwise), rounded to cents, so small usage reads"0.00". Mid-period figures come from the in-progress invoice and can still move. There is no per-request cost anywhere in the API.- Results are paginated: pass
limit(default 50, max 200) and follownext_cursor. - Rows break out cached-input, cache-write, and long-context tokens alongside input, output, search, and fetch, each with a billable twin.
Summary of what bills where
| Spend | Draws the allowance? | Bills wallet? |
|---|---|---|
Weighted token lanes on mindshub_air (allowance remaining) | yes | no |
Tokens on mindshub_air (allowance exhausted) | no | yes |
| Tokens on any other model | no | yes |
Cache writes on mindshub_air | yes, at weight 12 | only past the allowance |
Web search on mindshub_air | yes, 65,000 units per search | only past the allowance |
| Web search on any other model, and fetch | never | yes |
| Embeddings | never | yes |
| Jev decisions | never | no, priced $0; without credit or a topped-up card free of payment errors, capped daily (Jev decisions) |
count_tokens, GET /v1/models, refused requests | no | no |