Skip to main content

Getting started

MindsHub Inference gives you every major model behind one API key, reachable through the three APIs developers already use:

  • OpenAI Chat Completions: POST /v1/chat/completions
  • OpenAI Responses: POST /v1/responses
  • Anthropic Messages: POST /v1/messages

These are three request formats over the same engine, not three products. Streaming, tool calling, structured output, vision, and web search work through all of them, and any model works behind any of the three: Claude Fable from the OpenAI SDK, GPT 5.6 Sol from the Anthropic SDK. You keep the SDK and the code you already have, and change a base URL.

1. Get a key​

  1. Sign up at console.mindshub.ai.
  2. Create an API key.
  3. Copy it when it appears: the full key is shown once, at creation.
export MINDSHUB_API_KEY="mdb_..."

New organizations get a recurring weighted allowance on mindshub_air, so their first calls cost nothing. Paid requests draw a prepaid wallet. Jev decisions use a separate endpoint and are priced at $0 during the launch promotion. Until your organization has credit, or a card with a completed top-up and no payment error, Jev has a free daily allowance while shared free capacity lasts. See Billing.

The samples in these docs use mindshub_air wherever the choice of model is not the point, so you can run them before you add a card. Where a sample needs a different model, it says so at that sample.

2. The same request, three ways​

One prompt, sent through all three chat APIs. The Chat Completions tab uses mindshub_air, the model your included allowance covers. The other two name other vendors' models to show that any chat model works behind any chat API: Kimi K3 through the OpenAI SDK, GPT 5.6 Sol through the Anthropic SDK. Those two draw the prepaid wallet, so add funds first, or change the model to mindshub_air to use the allowance.

from openai import OpenAI
import os

client = OpenAI(
base_url="https://api.mindshub.ai/v1",
api_key=os.environ["MINDSHUB_API_KEY"],
)

response = client.chat.completions.create(
model="mindshub_air", # covered by your included allowance
messages=[{"role": "user", "content": "Name three uses for a paperclip."}],
)
print(response.choices[0].message.content)

Two details worth remembering:

  • The OpenAI SDKs take the base URL with /v1 and authenticate with api_key.
  • The Anthropic SDKs take the base URL without /v1 (the client appends the path itself) and authenticate with auth_token, not api_key. The api_key field sends an x-api-key header, which MindsHub rejects.

The TypeScript samples use top-level await, so run them as ES modules: a .mts/.mjs file, or "type": "module" in your package.json. Every other sample in these docs assumes a client built like one of the six above.

Not sure which API to use? Choosing an API has a one-screen decision table. Prefer curl? The one-minute version on the landing page has it.

3. Switch models by editing one string​

Every model is addressed by a short, stable alias. The alias resolves server-side to a concrete model at the provider, so upgrades don't break your code.

The five aliases in this loop all draw the prepaid wallet, because your included allowance covers mindshub_air only.

for model in ["fable", "gpt", "kimi", "deepseek-v4-flash", "gemini-flash"]:
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Explain a hash map in one sentence."}],
)
print(f"{model:14} {response.choices[0].message.content}")

That loop runs five models from five vendors on one key and one bill. The full list, with reasoning-effort support and live availability, is in Models and from GET /v1/models; what each model supports is on the capability matrix.

4. Go further​

Each guide shows one feature on all three APIs, with Python and TypeScript samples that run as pasted.

5. See what it cost​

Every non-streaming response carries usage; streams end with a usage chunk when you ask via stream_options: {"include_usage": true}:

print(response.usage.prompt_tokens, response.usage.completion_tokens)

And your account-wide totals are one call away:

curl "https://auth.mindshub.ai/v1/usage/summary/?range=period&group_by=model" \
-H "Authorization: Bearer $MINDSHUB_API_KEY"

That returns per-model token counts and cost for the current billing period, covering every SDK and coding agent you've pointed at MindsHub.

What we adapt for you​

Models disagree about parameters: some take no top_k, some take no reasoning_effort, and output ceilings differ. A parameter the target model can't take is dropped, a value above its range is clamped down, and the request is served either way, with every change named in the X-MindsHub-Dropped-Params and X-MindsHub-Clamped-Params response headers (except on a stream whose 200 went out early; see Errors). A model can still restrict the values of a parameter it supports; that error passes through as a 400 (Kimi K3 takes temperature only at 1). The full contract is in Core concepts.

Use it in your coding agent​

The same key runs your terminal and editor agents. Claude Code points at MindsHub with two environment variables and can run on Kimi K3 or any other catalog model, billed to the same balance as your application traffic. Codex runs on the same key and the same catalog.

Where to next​