OpenAI compatibility
Shardio implements a subset of the OpenAI dialect. The subset is deliberate, and the part that will bite you is not what we reject — it is what we accept and ignore.
Base URL and authentication
https://api.shardio.ai/v1
https://api.shardio.ai/openai/v1 is the same router under its canonical name; the bare
/v1 prefix is a permanent alias, so either is safe to hard-code. /anthropic/v1/* is
reserved and currently returns 404 with a "coming soon" message — it is not an
Anthropic-compatible endpoint yet.
Every endpoint except /v1/catalog/rates needs a bearer key:
Authorization: Bearer tsk_live_...
API keys authenticate the API and nothing else. They do not work against the console, and a console session does not work here. See API keys.
What we forward to the model
These six are validated, then passed to the engine. Anything you omit stays unset — we
never substitute a default of our own, so an omitted temperature is the engine's
default, not ours.
| Field | Type | Accepted range | Notes |
|---|---|---|---|
temperature | number | 0.0–2.0 | Outside the range is a 400, not a clamp. |
top_p | number | 0.0–1.0 | 0 is accepted; see below. |
seed | integer | signed 64-bit | Best-effort reproducibility, as everywhere. |
frequency_penalty | number | -2.0–2.0 | |
presence_penalty | number | -2.0–2.0 | |
stop | string or array | 1–4 entries, each 1–256 characters | An empty string would stop generation immediately, so it is refused. |
top_k is not accepted. It is not part of the dialect we advertise and its valid
range differs between engine versions, so it is ignored rather than translated.
top_p: 0
OpenAI accepts top_p: 0 and reads it as greedy decoding — take the single
highest-probability token. Some engines reject 0 outright. Because the whole product is
a base_url swap, a request that is valid against OpenAI must not fail here, so we
accept 0 and translate it to an epsilon (1e-6) that selects exactly that one token.
The effect is greedy decoding, which is what you asked for.
What we accept and ignore
Every field below is silently dropped. This list is not exhaustive — it is the set that real clients send most often — but the rule is: only the six fields above and the per-endpoint fields on Endpoints do anything.
| Field | What actually happens |
|---|---|
n | Ignored. You get exactly one choice, at index: 0, always. |
logit_bias | Ignored. No token biasing is applied. |
response_format | Ignored. There is no JSON mode; you get free-form text. |
tools, functions, tool_choice | Ignored. The model is never told tools exist and never emits a tool call. |
logprobs, top_logprobs | Ignored. No log probabilities are returned. |
stream_options | Ignored — including include_usage. A streamed response carries no usage chunk. |
user, metadata, store | Ignored. Attribute spend with a per-agent API key instead. |
top_k, min_p, repetition_penalty | Ignored. |
| Anything else | Ignored. |
cache_salt is the one field you cannot set even if you send it. The gateway owns it: it
partitions the prefix cache per account so no customer can observe another's cached
prefixes.
How chat messages reach the model
Your messages array reaches the engine intact: roles preserved, order preserved, one
turn per message. A system prompt is applied as a system prompt.
[
{ "role": "system", "content": "You are terse." },
{ "role": "user", "content": "Why is the sky blue?" }
]
arrives at the model as those two turns, rendered by the model's own chat template.
Supported roles are system, assistant, user and developer. An unrecognised role is a
400 naming the valid ones. tool is refused — it exists only to carry a tool result,
and we do not serve tool calling.
content
Either a string or an array of content parts:
{ "role": "user", "content": [{ "type": "text", "text": "Why is the sky blue?" }] }
Multiple text parts are joined with a newline, so the two forms are equivalent and cost
the same. Only text parts are accepted — an image_url or audio part is a 400
naming the part, because no model in this catalog takes other modalities and answering
from the text alone would be a confident answer to a question you did not ask.
What is refused rather than ignored
- Tool calling. A message carrying
tool_calls,tool_call_idorfunction_callis a400. We do not serve tools, and dropping the field would return a plausible completion to a request we had quietly mutilated. content: null. Send a string or a content-parts array.- An empty
messagesarray.
name is still ignored, as are unknown fields.
/v1/completions takes the same path: your prompt becomes one user message, so a raw
completion is still run through the model's chat template.
How tokens are counted
Gateway counts are the only billing truth. A node's self-reported usage never reaches money.
- Input is counted with the model's own tokenizer where one is pinned to that model, plus a flat +4 tokens per message as the structural allowance for a chat request.
- Models without a pinned tokenizer fall back to a byte approximation of about 4 bytes per token. It is deterministic but it is an approximation, and it can be off by tens of percent on unusual text.
- Output is always counted with the byte approximation, on every model, including those with a pinned tokenizer. Streaming detokenizes per frame, and a real tokenizer is not additive across frame boundaries — so the count that caps your spend and the count that bills you would disagree. Both use the same approximation instead, which keeps them identical.
The usage block in the response is exactly what was billed.
Limits and ceilings
| Limit | Default | Exceeded → |
|---|---|---|
| Request body | 10 MiB | 413 request_too_large |
| Requests per minute, per key | 600 | 429 rate_limit_requests |
| Tokens per minute, per key | 2,000,000 | 429 rate_limit_tokens |
| Concurrent streams, per key | 64 | 429 rate_limit_streams |
| Concurrent requests per model, across everyone | — | 429 rate_limit_sku_capacity |
Every 429 carries a Retry-After header in seconds. The rate limits are per key,
not per account, so splitting agents across keys splits their budgets too. The token
limit is measured after the fact — a request is refused when the current minute has
already crossed the line, so one large request can overshoot it.
Per-account overrides exist. Ask if you need one.