Credits

Shardio is prepaid. You buy credits, requests spend them, and there is no invoice at the end of the month. There are no plans and no tiers — credits are the only money, and everyone pays the same published rate.

The ledger holds integer microdollars: $1 = 1,000,000 µ. A single small request often costs less than a cent, so amounts are shown in microdollars where cents would round them to zero. 60 µ is $0.00006, and that is a real per-request price, not a rounding artefact.

Buying credits

Credit packs are $10, $50 and $200, and payment goes through a hosted checkout page.

Credits land when the payment completes, not when you are redirected back. A redirect you never finish moves no money, and a payment that clears credits your balance even if you close the tab. A top-up applies to your spendable balance immediately.

There is no free starter grant.

What a request costs

Input and output are priced separately, per 1M gateway-counted tokens:

cost = tokens_in × price_in + tokens_out × price_out

rounded to whole microdollars. Gateway counts are the only billing truth — the node that served the request does report its own usage, and that number never touches money.

The usage block in every response is exactly what you were charged for.

The reservation, and the 402 that surprises people

This is the part worth understanding before it costs you an afternoon.

Every request reserves its worst-case cost against your balance before it runs, and settles the difference afterwards. The reservation is sized from the request:

reserved = tokens_in × price_in + max_output × price_out

max_output = your max_tokens, if you sent one
             the model's full context limit, if you did not

Worked through on mid-14b (input $0.06 / 1M, output $0.15 / 1M), for a 1,000-token prompt:

With max_tokens: 500Without max_tokens
Reserved at admission135 µ4,975 µ
Actual cost (320 output tokens)108 µ108 µ
Refunded on completion27 µ4,867 µ

Same charge either way. Wildly different balance needed to be allowed to start.

Why it works this way

Without a reservation, a request's cost is only known once it has finished — by which point it has already been spent. Holding the worst case up front is what makes the promise "you cannot spend money you have not paid us" true for every individual request, including several running at once.

What happens at zero

Nothing queues. A request against a zero or negative balance is refused immediately with 402 insufficient_credits.

There are two distinct 402s, and they need different fixes:

codeMeansFix
insufficient_creditsYour balance is zero or below.Top up.
hold_exceeds_balanceYour balance is positive but smaller than this request's reservation. The response carries required_micro and available_micro.Lower max_tokens, or top up.

Neither is recorded in your request history — a refused request is not a request.

What you are and are not charged for

The rule underneath the table: a request is billed once at least one output token has reached you, and the charge is then the whole prompt plus every output token you received. The input is never prorated within a billed request — but a request that delivered nothing is not billed at all, prompt included.

OutcomeCharged
Completed normallyYes: prompt + output.
Hit the output cap (finish_reason: "length")Yes: prompt + what was generated.
Rate-limited (429)No. The reservation is released immediately.
Refused at admission (401, 402, 404)No.
No node available (503 no_capacity)No. Nothing ran.
Streamed request that failed or stalled after tokens reached youYes: prompt + the tokens that reached you.
Streamed request that failed before any token reached youNo.
You disconnected mid-streamYes: prompt + the tokens delivered before you disconnected.
Non-streamed request that failed (502)No — whether or not the node generated anything. The response has no content, so neither the output nor the prompt is charged.

The disconnect row is deliberate. Tokens you received cost a node real work, and gateway counts are billing truth — otherwise reading a full stream and hanging up before the end would be free inference.

If a request dies in a way that never settles at all — a process killed mid-flight — its reservation is reconciled against the recorded request afterwards and the unspent part is returned. You do not need to do anything, and you will not see the money disappear permanently.

Rate limits are separate

Running out of credits gives you a 402. Going too fast gives you a 429, and the two are independent: a healthy balance does not raise your rate limit, and a high rate limit does not let you spend money you do not have. Limits and codes are on OpenAI compatibility and Errors.