Credits
Shardio is prepaid. You buy credits, requests spend them, and there is no invoice at the end of the month. There are no plans and no tiers — credits are the only money, and everyone pays the same published rate.
The ledger holds integer microdollars: $1 = 1,000,000 µ. A single small request often
costs less than a cent, so amounts are shown in microdollars where cents would round them
to zero. 60 µ is $0.00006, and that is a real per-request price, not a rounding
artefact.
Buying credits
Credit packs are $10, $50 and $200, and payment goes through a hosted checkout page.
Credits land when the payment completes, not when you are redirected back. A redirect you never finish moves no money, and a payment that clears credits your balance even if you close the tab. A top-up applies to your spendable balance immediately.
There is no free starter grant.
What a request costs
Input and output are priced separately, per 1M gateway-counted tokens:
cost = tokens_in × price_in + tokens_out × price_out
rounded to whole microdollars. Gateway counts are the only billing truth — the node that served the request does report its own usage, and that number never touches money.
The usage block in every response is exactly what you were charged for.
The reservation, and the 402 that surprises people
This is the part worth understanding before it costs you an afternoon.
Every request reserves its worst-case cost against your balance before it runs, and settles the difference afterwards. The reservation is sized from the request:
reserved = tokens_in × price_in + max_output × price_out
max_output = your max_tokens, if you sent one
the model's full context limit, if you did not
Worked through on mid-14b (input $0.06 / 1M, output $0.15 / 1M), for a 1,000-token
prompt:
With max_tokens: 500 | Without max_tokens | |
|---|---|---|
| Reserved at admission | 135 µ | 4,975 µ |
| Actual cost (320 output tokens) | 108 µ | 108 µ |
| Refunded on completion | 27 µ | 4,867 µ |
Same charge either way. Wildly different balance needed to be allowed to start.
Why it works this way
Without a reservation, a request's cost is only known once it has finished — by which point it has already been spent. Holding the worst case up front is what makes the promise "you cannot spend money you have not paid us" true for every individual request, including several running at once.
What happens at zero
Nothing queues. A request against a zero or negative balance is refused immediately with
402 insufficient_credits.
There are two distinct 402s, and they need different fixes:
code | Means | Fix |
|---|---|---|
insufficient_credits | Your balance is zero or below. | Top up. |
hold_exceeds_balance | Your balance is positive but smaller than this request's reservation. The response carries required_micro and available_micro. | Lower max_tokens, or top up. |
Neither is recorded in your request history — a refused request is not a request.
What you are and are not charged for
The rule underneath the table: a request is billed once at least one output token has reached you, and the charge is then the whole prompt plus every output token you received. The input is never prorated within a billed request — but a request that delivered nothing is not billed at all, prompt included.
| Outcome | Charged |
|---|---|
| Completed normally | Yes: prompt + output. |
Hit the output cap (finish_reason: "length") | Yes: prompt + what was generated. |
Rate-limited (429) | No. The reservation is released immediately. |
Refused at admission (401, 402, 404) | No. |
No node available (503 no_capacity) | No. Nothing ran. |
| Streamed request that failed or stalled after tokens reached you | Yes: prompt + the tokens that reached you. |
| Streamed request that failed before any token reached you | No. |
| You disconnected mid-stream | Yes: prompt + the tokens delivered before you disconnected. |
Non-streamed request that failed (502) | No — whether or not the node generated anything. The response has no content, so neither the output nor the prompt is charged. |
The disconnect row is deliberate. Tokens you received cost a node real work, and gateway counts are billing truth — otherwise reading a full stream and hanging up before the end would be free inference.
If a request dies in a way that never settles at all — a process killed mid-flight — its reservation is reconciled against the recorded request afterwards and the unspent part is returned. You do not need to do anything, and you will not see the money disappear permanently.
Rate limits are separate
Running out of credits gives you a 402. Going too fast gives you a 429, and the two
are independent: a healthy balance does not raise your rate limit, and a high rate limit
does not let you spend money you do not have. Limits and codes are on
OpenAI compatibility and Errors.