Models
Put the model id from the first column into "model". Ids are stable: the engine,
quantisation and weights behind an id can change, the id you integrated against does not.
The catalogue
Prices are US dollars per 1M gateway-counted tokens, input and output priced separately.
| Model id | Class | Context | Input / 1M | Output / 1M | status |
|---|---|---|---|---|---|
tiny-3b | tiny | 8,192 | $0.02 | $0.04 | active |
small-8b | small | 16,384 | $0.03 | $0.07 | active |
small-7b-qwen | small | 32,768 | $0.04 | $0.08 | active |
mid-14b | mid | 32,768 | $0.06 | $0.15 | active |
reason-32b | reasoning | 32,768 | $0.10 | $0.30 | active |
embed-m3 | embed | 8,192 | $0.01 | — | active |
large-70b | large | 32,768 | — | — | coming |
nemotron35-30b | large | 1,048,576 | — | — | coming |
The status column is the literal string /v1/models and /v1/catalog/rates return, so a
client can branch on it. A coming model is listed but cannot be called: it has no price
and no build, and a request naming it returns 404 model_not_found. The same 404 answers a
model id that does not exist at all — we do not confirm which of the two it was.
embed-m3 has no output price because embeddings produce no output tokens. It is
reachable only from /v1/embeddings, and the generation endpoints refuse it — again with
404 model_not_found, so an embedding model in a chat call reads as a wrong model id
rather than as a mysterious failure.
What each one is for
| Model id | Base model | Quantisation | Use it for |
|---|---|---|---|
tiny-3b | Llama-3.2-3B-Instruct | 4-bit | Classification, routing, extraction — high-volume work where latency and price matter more than depth. |
small-8b | Llama-3.1-8B-Instruct | 4-bit | General chat and summarisation at a low rate. |
small-7b-qwen | Qwen2.5-7B-Instruct | 4-bit | The same class as small-8b with a 32k window and a different training mix; worth trying when small-8b is close but not right. |
mid-14b | Qwen2.5-14B-Instruct | 4-bit | Instruction following that the small class gets wrong. |
reason-32b | Qwen3-32B | 4-bit | Multi-step reasoning, where a longer, more expensive answer is the point. |
embed-m3 | BGE-M3 | fp16 | Retrieval and similarity. Embeddings only. |
nemotron35-30b | Nemotron 3.5 Lightning 30B | NVFP4 | Coming. A 1M-token context window. |
large-70b | 70B class | — | Coming. |
Reasoning models
A model whose build separates its thinking from its answer returns the thinking block in
choices[0].message.shardio_reasoning_content (and on the final delta of a stream), never
inside content — content is the answer. The field is absent entirely on a model that
does not reason, so nothing changes for the rest of the catalogue. Reasoning tokens are
billed as output: the model generated them, and usage.completion_tokens counts them.
GET /v1/catalog/rates carries reasoning per model and the same disclosure, so a cost
estimate can read it rather than assume.
Reading the live catalogue
Two endpoints answer this, and they differ in one way that matters:
GET /v1/catalog/ratesneeds no API key and is cacheable for 5 minutes. Use it for a pricing page or a cost estimate.GET /v1/modelsneeds a key, counts against your request-rate limit, and adds measuredttft_p50_msandtps_p50per model from the nodes currently serving it.
curl -s https://api.shardio.ai/v1/catalog/rates | jq '.data[] | {sku_id, price_in_micro_1m, price_out_micro_1m}'
Both report prices in microdollars per 1M tokens ($1 = 1,000,000 µ), because that is
the unit the ledger holds. price_in_micro_1m: 20000 is $0.02 per 1M input tokens.
Details of both responses: Endpoints.
What a context limit means here
The context limit is the model's window, and it is also the default size of the balance
reservation a request without max_tokens makes. On reason-32b that default is 32,768
output tokens — about 9,830 µ reserved for a request that may cost a hundredth of that.
The reservation is refunded, but it has to be affordable first. See
Credits.