Endpoints
Five endpoints, all under https://api.shardio.ai/v1. Anything else in the OpenAI
surface — files, batches, fine-tuning, assistants, images, audio, moderations — does not
exist here.
The decoding parameters (temperature, top_p, seed, the two penalties, stop) apply
to both generation endpoints and are documented once on
OpenAI compatibility.
POST /v1/chat/completions
| Field | Type | Required | Notes |
|---|---|---|---|
model | string | yes | A model id from Models. |
messages | array | yes | Objects with role and a content that is a string or an array of text parts. Roles reach the model — see OpenAI compatibility. |
max_tokens | integer ≥ 1 | no | Also sizes the balance reservation. Defaults to the model's context limit. |
stream | boolean | no | Defaults to false. |
{
"id": "chatcmpl-7kq3mNb9xZ4t",
"object": "chat.completion",
"created": 1787230800,
"model": "tiny-3b",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "…" },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 11, "completion_tokens": 8, "total_tokens": 19 }
}
choices always holds exactly one entry. finish_reason is "stop" or "length" and
nothing else — there is no "tool_calls" and no "content_filter". "length" means
generation hit its output cap, which is your max_tokens when you sent one and the
model's context limit when you did not.
There is no system_fingerprint, no service_tier, no logprobs, and no refusal
field. The id is derived from the request id we recorded, so
chatcmpl-7kq3mNb9xZ4t is request req_7kq3mNb9xZ4t in your usage history — quote it
when you ask us about a specific call.
POST /v1/completions
The legacy text-completion shape.
| Field | Type | Required | Notes |
|---|---|---|---|
model | string | yes | |
prompt | string or array of strings | no | An array is joined with newlines into one prompt. Defaults to "". |
max_tokens | integer ≥ 1 | no | As above. |
stream | boolean | no |
{
"id": "cmpl-7kq3mNb9xZ4t",
"object": "text_completion",
"created": 1787230800,
"model": "tiny-3b",
"choices": [{ "index": 0, "text": "…", "finish_reason": "stop" }],
"usage": { "prompt_tokens": 11, "completion_tokens": 8, "total_tokens": 19 }
}
An array prompt produces one completion over the joined text, not one completion per
element. echo, logprobs, best_of and suffix are ignored.
POST /v1/embeddings
| Field | Type | Required | Notes |
|---|---|---|---|
model | string | yes | Must be an embedding model. A generation model here is 404 model_not_found. |
input | string or array of strings | yes | One vector per element, in order. |
{
"object": "list",
"data": [{ "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456] }],
"model": "embed-m3",
"usage": { "prompt_tokens": 24, "total_tokens": 24 }
}
encoding_format and dimensions are ignored: vectors come back as float arrays at the
model's native width, never base64. Embeddings have no output tokens, so the reservation
and the charge are both input-only.
GET /v1/models
Requires a key, and counts against your requests-per-minute budget.
{
"object": "list",
"data": [
{
"id": "tiny-3b",
"object": "model",
"created": 1781049600,
"owned_by": "shardio",
"shardio": {
"class": "tiny",
"quant_label": "4-bit",
"context_limit": 8192,
"price_in_micro_1m": 20000,
"price_out_micro_1m": 40000,
"status": "active",
"ttft_p50_ms": 240,
"tps_p50": 61.4
}
}
]
}
shardio is a vendor extension: everything outside it is the OpenAI shape, and a client
that ignores unknown keys sees a standard /v1/models response.
ttft_p50_ms and tps_p50 are medians measured across the nodes currently able to serve
that model, and are null before any node has a baseline. Every listed model appears
here, including coming ones, which cannot be called at all. status: "active" means the
catalogue offers the model — not that a node is serving it today; see Models
for which ids can actually run.
GET /v1/catalog/rates
The price sheet. No key required, and served with Cache-Control: public, max-age=300
so a CDN can absorb it.
{
"object": "list",
"payout_share": 0.7,
"data": [
{
"sku_id": "tiny-3b",
"class": "tiny",
"quant_label": "4-bit",
"context_limit": 8192,
"price_in_micro_1m": 20000,
"price_out_micro_1m": 40000,
"status": "active"
}
]
}
payout_share is the fraction of every dollar that goes to the host who served the
request — 0.7. It is published because it is the same number for everyone, forever;
see Earnings.