Skip to main content

Billing

MindsHub Inference is pay-as-you-go, with no subscription. Two things fund your requests:

Included allowance. An organization with a free grant receives a recurring allowance for mindshub_air; one without a grant has an allowance of zero. It charges four token lanes at different weights, so efficient prompt caching preserves more useful work than uncached input or generated output. The window has a fixed duration and next_refresh_at gives its exact end. Programmatically, use included_percent_remaining; the old included_tokens object remains temporarily for compatibility and carries no token count despite its name.

The wallet. A prepaid, organization-level credit balance that funds paid models, embeddings, and web-search charges. Credits never expire. Jev is priced at $0 during its launch promotion; organizations without credit call it within a free daily allowance, described under Jev decisions.

A request passes the funding check when your wallet has available balance, when it uses mindshub_air with allowance remaining, or when it calls Jev and your organization has a card with a completed top-up and no payment error, or still has free Jev allowance left (Jev decisions). Authentication, access restrictions and rate limits still apply. Otherwise it is refused up front with 402 wallet_empty or 429 included_allowance_exhausted. Allowance exhaustion returns the same ISO-8601 reset_at in the JSON body and X-MindsHub-Reset-At header. An organization whose included allowance is zero (never granted, or set to zero) gets the same 429 with neither, because nothing refills; adding credit is the only way forward. Both billing 429s carry x-should-retry: false, so the OpenAI and Anthropic SDKs do not retry them. A refused request consumes nothing.

How the wallet works​

  • Top up from the console. Minimum top-up is $10.
  • Auto-recharge (optional): set a threshold and a target: when the balance crosses the threshold, your card is charged back up to the target, typically within a few minutes. A monthly recharge cap (default $100) limits automatic charges; hitting it pauses recharging until the calendar month rolls over.
  • Your card is only charged for top-ups and recharges. Usage consumes prepaid credits.
  • A declined auto-recharge stops recharging until the billing owner fixes the payment method; it never retries a declined card on its own. Requests keep working until the balance runs out.
  • Wallet visibility is role-based: the billing owner sees balances and dollar figures; other members see their own usage in tokens only.
  • An empty wallet never revokes API keys (Authentication); requests resume automatically once credit lands.

What gets metered​

Four token counters, per request:

CounterWhat it isBilled at
Input tokensPrompt tokens not served from cacheThe model's input rate
Output tokensEverything the model generates, including internal reasoningThe model's output rate
Cached input tokensPrompt tokens read from the provider's prompt cacheThe model's cached-input rate, roughly a tenth of input on most of the catalog but not uniformly so. Read it off the price list rather than deriving it
Cache-write tokensPrompt tokens written into the cacheThe model's cache-write rate: the input rate on most models, 1.25× input on the Claude family, mindshub_air, and the GPT 5.6 line. Not billed at all on the Gemini models or on mindshub_blaze, which shows n/a in the price list

On chat models, prompt caching is automatic and you never have to ask for it. On most of the catalog MindsHub marks the provider's cache breakpoints on the stable prefix of your request; on the rest the provider caches implicitly and there is nothing for us to mark. A request that already carries its own cache_control breakpoints, as Claude Code sends on Messages, is passed through exactly as you sent it. You see the effect as cached_tokens in responses and a lower effective input price. One rough edge: on Claude targets a system prompt sent as an array of content parts rather than a plain string gets no breakpoint of its own. The rolling breakpoint on the previous message picks it up from the second turn onward, so the cost is one extra full-price turn rather than every turn.

How this interacts with the included allowance:

  • On mindshub_air, cached reads draw 1 weighted unit per token, uncached input draws 10, cache writes draw 12, and output draws 60.
  • A web search on mindshub_air draws 65,000 weighted units and bills to the wallet only once the allowance is spent. On its own that is room for about 100 searches per window; the search results the model reads also draw the allowance as input tokens, so a real session fits fewer.
  • The weights are allowance policy, not prices. Wallet billing still uses the raw token counters and the current price list.
  • Page fetches, embeddings, and web searches on every other model always bill to the wallet, even when the request's model tokens were covered by the allowance.

The full matrix is the summary table at the bottom.

Price list​

These are MindsHub's prices; they track the upstream providers' list prices except for Jev, which is priced at $0 during its launch promotion. Per million tokens, with rates from September 2026 (the console is the authoritative, current source):

A fixed alias naming the same model as its moving alias — gpt-6-astra against gpt, sonnet-5 against sonnet — is priced identically to it and is left out of this table. Pinning a model costs nothing; it only stops the alias moving. Aliases for older releases price separately and are listed.

ModelInputOutputCached inputCache write
mindshub_air$0.20$1.20$0.02$0.25
mindshub_blaze$0.99$1.49$0.99n/a
gpt-oss-fireworks$0.15$0.60$0.02$0.15
haiku$1.00$5.00$0.10$1.25
sonnet$2.00$10.00$0.20$2.50
opus$5.00$25.00$0.50$6.25
fable$10.00$50.00$0.25$12.50
fable-5$10.00$50.00$1.00$12.50
gpt$10.00$50.00$1.00$12.50
gpt-5-6-sol$5.00$30.00$0.50$6.25
gpt-terra$2.00$12.00$0.20$2.50
gpt-luna$0.20$1.20$0.02$0.25
gpt-codex$1.75$14.00$0.18$1.75
gpt-mini$0.75$4.50$0.08$0.75
gpt-nano$0.20$1.25$0.02$0.20
gemini$2.00$12.00$0.20n/a
gemini-flash$0.75$3.75$0.08n/a
gemini-flash-3-7$0.75$3.75$0.08n/a
gemini-flash-3-6$0.75$3.75$0.08n/a
gemini-flash-3-5$1.50$9.00$0.15n/a
gemini-flash-3$0.50$3.00$0.05n/a
gemini-flash-lite$0.25$1.50$0.03n/a
kimi$3.00$15.00$0.30$3.00
deepseek-v4-flash$0.22$0.66$0.01$0.22
qwen$2.00$6.00$0.25$2.00
glm$1.40$4.40$0.26$1.40
glm-5-3-flash$0.15$0.50$0.03$0.15
muse-spark$1.25$4.25$0.15$1.25
muse-spark-1-2$1.25$4.25$0.15$1.25
muse-spark-1-1$1.25$4.25$0.15$1.25
grok$2.00$6.00$0.50$2.00
grok-4-6$2.00$6.00$0.50$2.00
grok-4-5$2.00$6.00$0.30$2.00
embed-small$0.02n/an/an/a
Jev decisions: jev, jev-1.13.0$0 (promotion)$0n/an/a

Aliases are described in Models. Jev also accepts jev-latest. Who can call Jev without a daily cap is under Jev decisions.

deepseek-v4-1-flash is the fixed alias for the model deepseek-v4-flash serves today, so it prices identically to it and is not listed separately above.

Jev decisions​

Jev input and output are priced at $0 during its launch promotion, and Jev requests never draw the included allowance. How much you can call it depends on how your organization is funded:

  • With wallet credit, or with a card that has completed at least one top-up and has no unresolved payment error, your organization calls Jev without a daily cap.
  • Otherwise, your organization gets a free daily allowance of 100,000 tokens and 100 requests, while shared free capacity lasts.

The allowance refills continuously rather than resetting once a day: it adds that amount back over every 24 hours and never holds more than one day's worth, so a new organization, which starts with a full allowance, can use roughly twice that on its first day. A call that starts while the allowance has requests and tokens left runs to completion even if it uses more tokens than remain, and the overrun is taken from the refill that follows. Every call that passes the funding check counts as one request, including one the provider refuses or times out.

A card on its own does not lift the cap; a completed top-up does. Past the allowance, or when shared free capacity runs out first, POST /v1/decisions returns 402 wallet_empty before the model runs. Without credit, the call works again once the allowance refills. The X-MindsHub-Recovery-Url header carries the console path for adding credit, and credit lifts the cap within about a minute of landing. Organization access restrictions and rate limits apply to every organization. Usage is measured and returned on every call, and decision requests have no cache charges.

Long-context pricing​

Some providers charge more for large prompts, and MindsHub passes those thresholds through. When the total prompt (uncached, cached, and cache-write tokens combined) exceeds the model's threshold, the higher rates apply to the entire request, on every dimension:

ModelThreshold (prompt tokens)Input / OutputCached inputCache write
mindshub_airabove 272,000$0.40 / $1.80$0.04$0.50
gptabove 272,000$20.00 / $75.00$2.00$25.00
gpt-5-6-solabove 272,000$10.00 / $45.00$1.00$12.50
gpt-terraabove 272,000$4.00 / $18.00$0.40$5.00
gpt-lunaabove 272,000$0.40 / $1.80$0.04$0.50
geminiabove 200,000$4.00 / $18.00$0.40n/a
grokat or above 200,000$4.00 / $12.00$1.00$4.00
grok-4-6at or above 200,000$4.00 / $12.00$1.00$4.00
grok-4-5at or above 200,000$4.00 / $12.00$0.60$4.00

Note the cached rates double too: a long, heavily-cached conversation is exactly the workload that crosses these thresholds. All other models price flat across their full context window.

Web search and fetch​

Per 1,000 uses, on models where web tools are available. mindshub_blaze, gpt-oss-fireworks and embed-small are not among them: web tools are dropped on those aliases, so no search or fetch charge can arise on them no matter what your tools array says. On mindshub_air a search draws the included allowance first and bills at the rate below only past it.

ModelsSearchFetch
mindshub_air, sonnet, opus, fable, fable-5, haiku, gpt, gpt-5-6-sol, gpt-terra, gpt-luna, gpt-codex, gpt-mini, gpt-nano$10.00n/a
kimi, deepseek-v4-flash, qwen, glm, glm-5-3-flash, muse-spark, muse-spark-1-2, muse-spark-1-1$7.00$1.00
gemini, gemini-flash, gemini-flash-3-7, gemini-flash-3-6, gemini-flash-3-5, gemini-flash-3, gemini-flash-lite$14.00n/a
grok, grok-4-6, grok-4-5$5.00n/a

Where no fetch price is listed, fetched pages bill as ordinary input tokens rather than per fetch.

Remote MCP​

No per-call charge. The MCP server is yours and neither upstream charges to call it; what you pay is the tokens the server's tool schemas and results add to the turn, which usage reports as ordinary input and output tokens. Calls are counted (mcp_call) on your usage summary so the count is visible; the count carries no price today. See Remote MCP servers.

Checking usage and balance from code​

Two account endpoints accept your API key. They live on auth.mindshub.ai (not api.mindshub.ai), and the trailing slash is part of the path.

GET https://auth.mindshub.ai/v1/entitlements/me/: what you can use right now:

{
"wallet": {"balance_usd": "90.45", "can_consume": true, "auto_recharge": {"enabled": true, "...": "..."}},
"is_billing_owner": true,
"models": [
{"id": "mindshub_air", "label": "MindsHub Air", "free_bucket": true, "embedding": false,
"enabled": true, "reasoning_efforts": null, "default_reasoning_effort": "medium"}
],
"included_percent_remaining": 42.5,
"next_refresh_at": "2026-08-25T18:20:26Z"
}

next_refresh_at is when the allowance refills. free_bucket marks the model it covers. included_percent_remaining is canonical and keeps fractional values, so 0.4% does not display as exhausted. The response also carries a deprecated included_tokens object for older clients, omitted from the example above and scheduled for removal. On a live allowance it restates the same proportion out of a limit of 100 and says so in a unit field; an uncapped allowance reports every counter as null, and an organization with no grant reports a limit of 0. Read included_percent_remaining instead. The absolute size of the allowance is not published.

GET https://auth.mindshub.ai/v1/usage/summary/?range=period&group_by=model: what you've consumed. Response (excerpt; some fields omitted for brevity):

{
"range": {"requested": "period", "start": "2026-07-01T00:00:00Z", "end": "2026-08-01T00:00:00Z"},
"results": [
{"dimensions": {"model_alias": "opus"},
"usage": {"input_tokens": 218809, "output_tokens": 226135,
"billable_input_tokens": 218809, "billable_output_tokens": 226135,
"search_count": 87, "fetch_count": 16},
"cost_usd": "19.54"}
],
"next_cursor": "eyJvIjogM30=",
"totals": {"usage": {"...": "..."}, "cost_usd": "19.54"}
}

Reading it correctly:

  • This endpoint reports paid consumption only. range=period is the monthly paid billing period. The allowance runs on a separate fixed-duration window and is not reported here at all. Read it from GET /v1/entitlements/me/, where included_percent_remaining gives the proportion left and next_refresh_at gives the instant the window refills.
  • cost_usd is a decimal string, populated only for the billing owner (null otherwise), rounded to cents, so small usage reads "0.00". Mid-period figures come from the in-progress invoice and can still move. There is no per-request cost anywhere in the API.
  • Results are paginated: pass limit (default 50, max 200) and follow next_cursor.
  • Rows break out cached-input, cache-write, and long-context tokens alongside input, output, search, and fetch, each with a billable twin.

Summary of what bills where​

SpendDraws the allowance?Bills wallet?
Weighted token lanes on mindshub_air (allowance remaining)yesno
Tokens on mindshub_air (allowance exhausted)noyes
Tokens on any other modelnoyes
Cache writes on mindshub_airyes, at weight 12only past the allowance
Web search on mindshub_airyes, 65,000 units per searchonly past the allowance
Web search on any other model, and fetchneveryes
Embeddingsneveryes
Jev decisionsneverno, priced $0; without credit or a topped-up card free of payment errors, capped daily (Jev decisions)
count_tokens, GET /v1/models, refused requestsnono