Endpoints

Five endpoints, all under https://api.shardio.ai/v1. Anything else in the OpenAI surface — files, batches, fine-tuning, assistants, images, audio, moderations — does not exist here.

The decoding parameters (temperature, top_p, seed, the two penalties, stop) apply to both generation endpoints and are documented once on OpenAI compatibility.

POST /v1/chat/completions

FieldTypeRequiredNotes
modelstringyesA model id from Models.
messagesarrayyesObjects with role and a content that is a string or an array of text parts. Roles reach the model — see OpenAI compatibility.
max_tokensinteger ≥ 1noAlso sizes the balance reservation. Defaults to the model's context limit.
streambooleannoDefaults to false.
{
  "id": "chatcmpl-7kq3mNb9xZ4t",
  "object": "chat.completion",
  "created": 1787230800,
  "model": "tiny-3b",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "…" },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 11, "completion_tokens": 8, "total_tokens": 19 }
}

choices always holds exactly one entry. finish_reason is "stop" or "length" and nothing else — there is no "tool_calls" and no "content_filter". "length" means generation hit its output cap, which is your max_tokens when you sent one and the model's context limit when you did not.

There is no system_fingerprint, no service_tier, no logprobs, and no refusal field. The id is derived from the request id we recorded, so chatcmpl-7kq3mNb9xZ4t is request req_7kq3mNb9xZ4t in your usage history — quote it when you ask us about a specific call.

POST /v1/completions

The legacy text-completion shape.

FieldTypeRequiredNotes
modelstringyes
promptstring or array of stringsnoAn array is joined with newlines into one prompt. Defaults to "".
max_tokensinteger ≥ 1noAs above.
streambooleanno
{
  "id": "cmpl-7kq3mNb9xZ4t",
  "object": "text_completion",
  "created": 1787230800,
  "model": "tiny-3b",
  "choices": [{ "index": 0, "text": "…", "finish_reason": "stop" }],
  "usage": { "prompt_tokens": 11, "completion_tokens": 8, "total_tokens": 19 }
}

An array prompt produces one completion over the joined text, not one completion per element. echo, logprobs, best_of and suffix are ignored.

POST /v1/embeddings

FieldTypeRequiredNotes
modelstringyesMust be an embedding model. A generation model here is 404 model_not_found.
inputstring or array of stringsyesOne vector per element, in order.
{
  "object": "list",
  "data": [{ "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456] }],
  "model": "embed-m3",
  "usage": { "prompt_tokens": 24, "total_tokens": 24 }
}

encoding_format and dimensions are ignored: vectors come back as float arrays at the model's native width, never base64. Embeddings have no output tokens, so the reservation and the charge are both input-only.

GET /v1/models

Requires a key, and counts against your requests-per-minute budget.

{
  "object": "list",
  "data": [
    {
      "id": "tiny-3b",
      "object": "model",
      "created": 1781049600,
      "owned_by": "shardio",
      "shardio": {
        "class": "tiny",
        "quant_label": "4-bit",
        "context_limit": 8192,
        "price_in_micro_1m": 20000,
        "price_out_micro_1m": 40000,
        "status": "active",
        "ttft_p50_ms": 240,
        "tps_p50": 61.4
      }
    }
  ]
}

shardio is a vendor extension: everything outside it is the OpenAI shape, and a client that ignores unknown keys sees a standard /v1/models response.

ttft_p50_ms and tps_p50 are medians measured across the nodes currently able to serve that model, and are null before any node has a baseline. Every listed model appears here, including coming ones, which cannot be called at all. status: "active" means the catalogue offers the model — not that a node is serving it today; see Models for which ids can actually run.

GET /v1/catalog/rates

The price sheet. No key required, and served with Cache-Control: public, max-age=300 so a CDN can absorb it.

{
  "object": "list",
  "payout_share": 0.7,
  "data": [
    {
      "sku_id": "tiny-3b",
      "class": "tiny",
      "quant_label": "4-bit",
      "context_limit": 8192,
      "price_in_micro_1m": 20000,
      "price_out_micro_1m": 40000,
      "status": "active"
    }
  ]
}

payout_share is the fraction of every dollar that goes to the host who served the request — 0.7. It is published because it is the same number for everyone, forever; see Earnings.