Quickstart
Shardio speaks the OpenAI dialect. Change base_url, use a Shardio key, keep the rest
of your code.
1. Get a key
A key looks like tsk_live_<reference>_<secret>. You are handed the secret once — it is
stored hashed, so nobody, including us, can read it back afterwards. Name it after the
agent that will use it; that name is the only spend attribution you get. See
API keys.
Keys are prepaid: your account needs credits before the first call, because a zero balance
returns 402 immediately rather than queueing. Credits are not self-serve yet either. See
Credits.
2. Call it
curl https://api.shardio.ai/v1/chat/completions \
-H "Authorization: Bearer $SHARDIO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tiny-3b",
"messages": [{"role": "user", "content": "Say hello in five words."}],
"max_tokens": 32
}'import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.shardio.ai/v1",
api_key=os.environ["SHARDIO_API_KEY"],
)
response = client.chat.completions.create(
model="tiny-3b",
messages=[{"role": "user", "content": "Say hello in five words."}],
max_tokens=32,
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.shardio.ai/v1",
apiKey: process.env.SHARDIO_API_KEY,
});
const response = await client.chat.completions.create({
model: "tiny-3b",
messages: [{ role: "user", content: "Say hello in five words." }],
max_tokens: 32,
});
console.log(response.choices[0].message.content);That is the entire integration. No Shardio SDK to install, no request shape to learn, no proxy to run — the OpenAI client you already use, pointed somewhere else.
3. Read the response
{
"id": "chatcmpl-7kq3mNb9xZ4t",
"object": "chat.completion",
"created": 1787230800,
"model": "tiny-3b",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello there, good to meet you." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 11, "completion_tokens": 8, "total_tokens": 19 }
}
usage is what you are billed on. Those counts come from the gateway, never from the
node that served the request — a node's own numbers never reach money.
Always send max_tokens
Streaming
Pass stream: true for server-sent events in the OpenAI chunk format.
stream = client.chat.completions.create(
model="tiny-3b",
messages=[{"role": "user", "content": "Count to ten."}],
max_tokens=64,
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Streaming has one wrinkle worth knowing before you ship: because the response commits to
200 before the first token, a mid-stream failure arrives as an error event, not an
error status. See Streaming.
Next
- Models — what to put in
model, and what each costs. - OpenAI compatibility — which parameters we forward, and which we ignore.
- Errors — every code, its cause and its fix.