Errors
Every error uses the OpenAI envelope. code is the machine-readable part — branch on it,
not on the message text, which we reserve the right to improve.
{
"error": {
"message": "Insufficient credits. Top up to continue.",
"type": "insufficient_credits",
"code": "insufficient_credits"
}
}
Some errors add fields alongside message, type and code. hold_exceeds_balance is
the one that does today.
API errors
These are the codes a client calling https://api.shardio.ai/v1 can receive.
| Status | code | Cause | Fix |
|---|---|---|---|
| 400 | (none) | The body is not valid JSON, or a field failed validation. | The message names the field and the constraint: temperature: Input should be less than or equal to 2. Fix that field. |
| 401 | invalid_api_key | No Authorization header, a header that is not Bearer, a malformed key, an unknown key, or a revoked key. All four answer identically. | Check the header is Authorization: Bearer tsk_live_… and the key is not revoked. A key revoked in the last few minutes may still work, and a token that was rejected is remembered as bad for up to a minute — see API keys. |
| 402 | insufficient_credits | Your balance is zero or negative. Nothing queues at zero. | Buy credits. See Credits. |
| 402 | hold_exceeds_balance | Your balance is positive, but smaller than the reservation this request needs. Carries required_micro and available_micro. | Send a smaller max_tokens, or top up. This is the error a caller with a healthy-looking balance hits after omitting max_tokens on a long-context model. |
| 404 | model_not_found | The model id does not exist, is not yet available, or is the wrong kind for this endpoint (an embedding model on /chat/completions, or a chat model on /embeddings). | Check GET /v1/models and the status field. |
| 404 | (none), type not_found_error | You called /anthropic/v1/…. | That dialect is reserved and not implemented. Use the OpenAI dialect. |
| 413 | request_too_large | The request body exceeds 10 MiB. Checked before the body is buffered, so a lying Content-Length does not get around it. | Send less. If a single prompt is genuinely this large, it exceeds every model's context window anyway. |
| 429 | rate_limit_requests | More than 600 requests in the current minute, on this key. | Back off for the Retry-After seconds. Limits are per key, so a noisy agent only throttles itself. |
| 429 | rate_limit_tokens | This key's token budget for the current minute is already spent (2,000,000 by default). Measured after the fact, so one large request can overshoot and refuse the next. | Wait out Retry-After, or spread work across keys. |
| 429 | rate_limit_streams | More than 64 concurrent in-flight requests on this key. | Reduce concurrency, or use more keys. |
| 429 | rate_limit_sku_capacity | The whole platform is at its admission ceiling for that model — not your key's fault. | Retry after 1 second, or use a different model. |
| 502 | upstream_failed | A node accepted the request and reported that it did not finish. Never retried by us: the generation already ran once. | Retry the request. Whatever the node generated before failing is billed, along with the prompt — see Credits, because on a non-streamed request none of it is in the response. |
| 502 | upstream_stall | The token stream went quiet for 2 seconds. Terminal for one of two reasons: output had already reached you, so re-issuing would duplicate it; or nothing had, and the one permitted retry has already been spent. | Retry. If the node generated output, that output and the prompt are billed; if it generated none, nothing was. |
| 503 | no_capacity | No node was available to take the request before the dispatch deadline. Nothing ran. | Retry shortly. Nothing was billed. This is the code to expect for a model with no node currently serving it. |
A 429 always carries a Retry-After header, in seconds. A 429 is also recorded in
your request history, so a throttled minute is visible after the fact; a 401 or 402
is not recorded.
The three server-side codes, and how they differ
They look similar and mean different things, which matters when you decide whether to retry:
no_capacity(503) — no node was ever reached. Zero cost, always safe to retry.upstream_failed(502) — a node ran the request and failed it. Safe to retry, but it costs a second generation, and the first is billed only for output that reached you.upstream_stall(502) — a node went quiet. Usually you hold partial output and a retry produces a whole second answer; it also arrives with no output at all when neither the first attempt nor the retry produced a single token. In that case nothing reached the gateway and nothing was billed.
On a streamed request none of these can be an HTTP status, because the response already
committed to 200. They arrive as a terminal error event instead —
see Streaming.
Console errors
These come from the console API, which authenticates with a signed-in session rather than an API key. There is no console site in front of it yet, so you will only see these if you are calling that API directly.
| Status | code | Cause |
|---|---|---|
| 401 | not_authenticated | No valid session. Sign in. |
| 403 | no_account | The identity is signed in but has no Shardio account. |
| 403 | account_suspended | The identity's accounts are all suspended. Distinct from no_account on purpose: we do not tell a suspended user they never existed. |
| 403 | member_only | The endpoint acts on a member account and this session has none (an operator). |
| 403 | cross_site_blocked | A state-changing console request arrived cross-site. |
| 400 | invalid_cursor | A pagination cursor was malformed. Start the page again. |
| 422 | unknown_model | A filter named a model that is not in the catalogue. |
| 404 | not_found | No such request id in this account's history. |
Buying credits
| Status | code | Cause | Fix |
|---|---|---|---|
| 404 | billing_disabled | Credit purchase is not enabled on this deployment. The route behaves as if it does not exist. | Nothing you can do client-side. |
| 403 | not_a_builder | A host account tried to buy credits. Credits are builder money: a host account has no keys to spend them with. | Switch to your builder account. |
| 422 | unknown_pack | The pack_id is not one of the offered packs. | Read the pack list and use an id from it. |
| 502 | stripe_unavailable | Our payment provider refused to start a checkout session. | Try again. Nothing was charged. |
| 400 | invalid_signature | A payment webhook failed signature verification. Provider-side only. | — |
Node errors
These come from the /agent API that the shardio CLI uses. They surface as CLI output
rather than as raw JSON; see Join the mesh.
| Status | code | Cause |
|---|---|---|
| 401 | invalid_host_token | Missing, malformed, unknown or revoked host token. Run shardio login again. |
| 404 | node_not_found | No such node, or the node belongs to another account. Deliberately the same answer, so nobody can enumerate the mesh. |
| 409 | node_quarantined | The node is quarantined. It cannot resume or mint credentials until it is released. |
| 409 | no_elected_builds | The node asked for worker credentials before electing any model. Run shardio setup --build … first. |
| 501 | not_implemented | The platform has no dispatch plane configured, so there is nothing to mint credentials against. Platform-side. |
| 502 | flux_admin_error | The dispatch plane refused to mint a join token. Retry; if it persists it is ours to fix. |