Streaming

Send "stream": true and the response is text/event-stream in the OpenAI chunk format. Any OpenAI client reads it without modification.

The chunk sequence

A chat stream opens with a role chunk carrying no content, then one chunk per delta, then a chunk carrying finish_reason and an empty delta, then the terminator.

data: {"id":"chatcmpl-7kq3mNb9xZ4t","object":"chat.completion.chunk","created":1787230800,"model":"tiny-3b","choices":[{"index":0,"finish_reason":null,"delta":{"role":"assistant","content":""}}]}

data: {"id":"chatcmpl-7kq3mNb9xZ4t","object":"chat.completion.chunk","created":1787230800,"model":"tiny-3b","choices":[{"index":0,"finish_reason":null,"delta":{"content":"Hello"}}]}

data: {"id":"chatcmpl-7kq3mNb9xZ4t","object":"chat.completion.chunk","created":1787230800,"model":"tiny-3b","choices":[{"index":0,"finish_reason":"stop","delta":{}}]}

data: [DONE]

/v1/completions streams the same sequence with "object":"text_completion" and a "text" field in place of delta, and no opening role chunk.

There is no usage chunk. stream_options: {"include_usage": true} is ignored, so a streamed request never reports its token counts inline. They are recorded against the request id embedded in id, and that per-request detail is kept for 90 days; account totals are permanent. If you need the counts for longer than that, keep them yourself.

How a stream fails

The response commits to 200 before the model produces anything, so a failure after that point cannot be an HTTP status. It arrives as a terminal error event instead:

data: {"error":{"type":"server_error","code":"upstream_stall","message":"Upstream token stream stalled."}}

When that happens there is no finish_reason chunk and no [DONE]. That absence is the signal — a stream that ends without [DONE] did not finish.

Three codes can arrive this way. Each has the same meaning as its non-streaming HTTP counterpart, described on Errors:

CodeMeaning
no_capacityNo node took the request. Nothing was generated and nothing was billed.
upstream_failedThe serving node started and did not finish. The tokens that reached you are billed, with the whole prompt.
upstream_stallThe stream went quiet for 2 seconds. The tokens that reached you are billed, with the whole prompt; if none reached you, nothing is billed.

Retries and duplicate output

We retry at most once, and only before the first token reaches you. Once you have received a single delta, a stall is terminal: re-issuing would duplicate output you already have. A stall on the retry is terminal too, even though you hold nothing — the one retry is spent, so upstream_stall can arrive with an empty stream. Nothing is billed when it does.

A node that positively reports failure is never retried at all, even before the first token. It already ran the generation; a retry buys a second one rather than a second chance.

A stream that produces no token for 2 seconds is considered dead. That is a gap between tokens, not a total time limit — a slow-but-alive generation runs as long as it needs.

Cancelling

Close the connection. The request is recorded as cancelled and the upstream generation is stopped rather than left running.

Tokens already delivered are still billed, and so is the whole prompt. Gateway counts are billing truth, and reading a full stream then disconnecting before [DONE] would otherwise be free inference.

The output cap still applies

A stream stops at the output budget your reservation funded — your max_tokens, or the model's context limit if you sent none — and the final chunk reads "finish_reason": "length". See Credits.