Models

Put the model id from the first column into "model". Ids are stable: the engine, quantisation and weights behind an id can change, the id you integrated against does not.

The catalogue

Prices are US dollars per 1M gateway-counted tokens, input and output priced separately.

Model idClassContextInput / 1MOutput / 1Mstatus
tiny-3btiny8,192$0.02$0.04active
small-8bsmall16,384$0.03$0.07active
small-7b-qwensmall32,768$0.04$0.08active
mid-14bmid32,768$0.06$0.15active
reason-32breasoning32,768$0.10$0.30active
embed-m3embed8,192$0.01—active
large-70blarge32,768——coming
nemotron35-30blarge1,048,576——coming

The status column is the literal string /v1/models and /v1/catalog/rates return, so a client can branch on it. A coming model is listed but cannot be called: it has no price and no build, and a request naming it returns 404 model_not_found. The same 404 answers a model id that does not exist at all — we do not confirm which of the two it was.

embed-m3 has no output price because embeddings produce no output tokens. It is reachable only from /v1/embeddings, and the generation endpoints refuse it — again with 404 model_not_found, so an embedding model in a chat call reads as a wrong model id rather than as a mysterious failure.

What each one is for

Model idBase modelQuantisationUse it for
tiny-3bLlama-3.2-3B-Instruct4-bitClassification, routing, extraction — high-volume work where latency and price matter more than depth.
small-8bLlama-3.1-8B-Instruct4-bitGeneral chat and summarisation at a low rate.
small-7b-qwenQwen2.5-7B-Instruct4-bitThe same class as small-8b with a 32k window and a different training mix; worth trying when small-8b is close but not right.
mid-14bQwen2.5-14B-Instruct4-bitInstruction following that the small class gets wrong.
reason-32bQwen3-32B4-bitMulti-step reasoning, where a longer, more expensive answer is the point.
embed-m3BGE-M3fp16Retrieval and similarity. Embeddings only.
nemotron35-30bNemotron 3.5 Lightning 30BNVFP4Coming. A 1M-token context window.
large-70b70B class—Coming.

Reasoning models

A model whose build separates its thinking from its answer returns the thinking block in choices[0].message.shardio_reasoning_content (and on the final delta of a stream), never inside content — content is the answer. The field is absent entirely on a model that does not reason, so nothing changes for the rest of the catalogue. Reasoning tokens are billed as output: the model generated them, and usage.completion_tokens counts them. GET /v1/catalog/rates carries reasoning per model and the same disclosure, so a cost estimate can read it rather than assume.

Reading the live catalogue

Two endpoints answer this, and they differ in one way that matters:

  • GET /v1/catalog/rates needs no API key and is cacheable for 5 minutes. Use it for a pricing page or a cost estimate.
  • GET /v1/models needs a key, counts against your request-rate limit, and adds measured ttft_p50_ms and tps_p50 per model from the nodes currently serving it.
curl -s https://api.shardio.ai/v1/catalog/rates | jq '.data[] | {sku_id, price_in_micro_1m, price_out_micro_1m}'

Both report prices in microdollars per 1M tokens ($1 = 1,000,000 µ), because that is the unit the ledger holds. price_in_micro_1m: 20000 is $0.02 per 1M input tokens.

Details of both responses: Endpoints.

What a context limit means here

The context limit is the model's window, and it is also the default size of the balance reservation a request without max_tokens makes. On reason-32b that default is 32,768 output tokens — about 9,830 µ reserved for a request that may cost a hundredth of that. The reservation is refunded, but it has to be affordable first. See Credits.