Join the mesh
Shardio does not own GPUs. Requests are served by machines like yours, and the host who serves a request earns 70% of what the customer paid for it. The share is the same for every node and every model; see Earnings.
What you need
| Requirement | |
|---|---|
| GPU | An NVIDIA card with a working driver — nvidia-smi must run. |
| VRAM | Enough for the models you elect. See the table below. |
| OS | Linux. |
| Container runtime | docker or podman. Rootless podman is preferred. |
| Core dumps | kernel.core_pattern must not start with |. A piped core collector would copy an engine's whole address space — including prompts — to a third-party handler. |
You do not need nvidia-container-toolkit. The per-request sealed container is
CPU-only; it reaches the GPU through the engine, never directly.
VRAM floor per model, from the catalogue:
| Model | Minimum VRAM |
|---|---|
tiny-3b | 2 GB |
embed-m3 | 3 GB |
small-8b | 7 GB |
small-7b-qwen | 7 GB |
mid-14b | 11 GB |
reason-32b | 22 GB |
Electing several models on one card splits its memory between them evenly, with 10% held back as headroom.
Nothing checks that the split fits your card. The supervisor divides the card by the
number of elected builds, and the only thing it refuses is a share below 10% — ten or more
builds, whatever the card is. Elect mid-14b and reason-32b on a 16 GiB card and each
engine is handed 45% of it, about 7 GiB, under either model's floor: nothing refuses, and
the engine fails when it loads. Check the floors above against the card you actually have
before you elect.
The four commands
shardio check
Run it before anything else. It probes the machine and prints one line per check, then
host is ready or host is NOT ready, exiting non-zero if a hard check failed.
Each line reads [ok |warn|FAIL] <check>: <detail>, and the last line is the verdict:
$ shardio check
[ok ] gpu: …
[ok ] runtime: …
[ok ] userns: …
[ok ] bind_mounts: …
[warn] core_pattern: …
host is ready
Five checks, in that order. Three are hard — a GPU, a container runtime, and user
namespaces. Two are warnings: bind_mounts, which fails on Docker Desktop under WSL2
because it cannot bind-mount the paths a node needs and fails silently rather than
loudly, and core_pattern, which fails on a piped core collector. --json prints the
same result as machine-readable output.
shardio login
shardio login
You are not asked for an endpoint — there is one mesh, and the CLI already knows where it
is. (--api-url exists for pointing a node at a local gateway during development; a host
never needs it.)
The token can come from SHARDIO_HOST_TOKEN instead, or be typed at a hidden prompt. The
CLI verifies it against the platform before writing anything, then stores it at
~/.config/shardio/config.toml with mode 0600.
shardio setup
shardio setup --name lakehouse-server
Registers the node and takes a snapshot of what it has — GPU model, VRAM, cores, RAM.
With no --build, it prints the catalogue with a fit mark against each model and stops:
✓ tiny-3b min_vram_gb=2
✓ small-8b min_vram_gb=7
✗ reason-32b min_vram_gb=22
re-run with --build <id> to elect.
Then elect what you want to serve. --build repeats:
shardio setup --name lakehouse-server --build bld_tiny-3b-vllm-r1
Electing writes the manifest your node will run: which engine, which weights, and the digest those weights must hash to.
The capability snapshot is advisory. It decides what the CLI offers you; it is not evidence, and the platform never treats it as proof of anything.
shardio up
shardio up --detach
This is the long one. It fetches and verifies weights, loads the engine, waits for the engine to report healthy, and only then asks the platform for the credentials that register the node for work. On a cold machine, expect several minutes — the default wait is 30 minutes.
Credentials are minted last, deliberately: they are one-time join tokens with a short life, so minting them before the engine is warm would burn them waiting.
The rest
| Command | Does |
|---|---|
shardio status | Local process state cross-checked against the platform's view, per elected build. --json available. |
shardio logs [-f] [--build <id>] | Supervisor and per-build engine logs. |
shardio down | Stops everything, and reaps orphaned engines and workers. --grace sets the SIGTERM window, default 30s. |
shardio pause / shardio resume | Records the node as paused or serving in the registry, and nothing more — read the note below before you rely on it. |
shardio config path|get|set | Shows where config lives, prints it with the token masked, or sets api_url. |
What actually runs on your machine
Three layers, and the distinction matters because only one of them is exposed to customer input.
The supervisor is one process you start. It fetches weights, starts engines, and starts one dispatch worker per elected model. It holds your credentials.
One engine per elected model, warm and resident, reachable only on the shard's own network. This is the only thing that touches your GPU.
One sealed container per request, started and destroyed for each individual request. It has:
- no GPU access — it reaches the engine here over a local HTTP connection;
- no credentials — it has never held your host token and cannot ask for one;
- no access to your files — nothing of yours is mounted into it;
- a hard memory, CPU, process and lifetime cap;
- a read-only, single-purpose program: read the request, stream the answer back.
Prompts and completions do pass through that connection to your engine, because that is what inference is. They are not written to your disk by anything Shardio runs, and because the container is created for one request and destroyed with it, nothing from one request is left behind for the next.
Weights are verified, not trusted
Every model artefact is pinned to a sha256 in the catalogue. Your node downloads it, hashes it, checks its length, and only then moves it into place. A mismatch is deleted and the build refuses to serve — a corrupted or substituted artefact never gets loaded, and never certifies.
After shardio up
Your elected builds start in certifying. Being able to serve a model and being allowed
to serve it are separate facts, and the second one is ours to establish, not yours. See
Certification.