Quickstart

Shardio speaks the OpenAI dialect. Change base_url, use a Shardio key, keep the rest of your code.

1. Get a key

A key looks like tsk_live_<reference>_<secret>. You are handed the secret once — it is stored hashed, so nobody, including us, can read it back afterwards. Name it after the agent that will use it; that name is the only spend attribution you get. See API keys.

Keys are prepaid: your account needs credits before the first call, because a zero balance returns 402 immediately rather than queueing. Credits are not self-serve yet either. See Credits.

2. Call it

curl https://api.shardio.ai/v1/chat/completions \
-H "Authorization: Bearer $SHARDIO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "model": "tiny-3b",
  "messages": [{"role": "user", "content": "Say hello in five words."}],
  "max_tokens": 32
}'

That is the entire integration. No Shardio SDK to install, no request shape to learn, no proxy to run — the OpenAI client you already use, pointed somewhere else.

3. Read the response

{
  "id": "chatcmpl-7kq3mNb9xZ4t",
  "object": "chat.completion",
  "created": 1787230800,
  "model": "tiny-3b",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hello there, good to meet you." },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 11, "completion_tokens": 8, "total_tokens": 19 }
}

usage is what you are billed on. Those counts come from the gateway, never from the node that served the request — a node's own numbers never reach money.

Always send max_tokens

Streaming

Pass stream: true for server-sent events in the OpenAI chunk format.

stream = client.chat.completions.create(
    model="tiny-3b",
    messages=[{"role": "user", "content": "Count to ten."}],
    max_tokens=64,
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Streaming has one wrinkle worth knowing before you ship: because the response commits to 200 before the first token, a mid-stream failure arrives as an error event, not an error status. See Streaming.

Next

  • Models — what to put in model, and what each costs.
  • OpenAI compatibility — which parameters we forward, and which we ignore.
  • Errors — every code, its cause and its fix.