Cluster and Hardware Commands

fdl probe

Cluster readiness audit. Single-host (default) probes the local box; in cluster context (fdl @cluster probe or FDL_ENV=cluster fdl probe) SSHes every host in fdl.cluster.yml and aggregates the report. Use it as a CI gate before launching multi-hour training runs.

fdl probe                       # local host: GPU + libtorch + NCCL + shared-data
fdl @cluster probe               # multi-host: SSHes each worker, aggregates
fdl probe --json                # machine-readable for CI gating
fdl @cluster probe --json        # cluster JSON aggregate
fdl probe --skip-mount          # skip shared-data-mount check on single-host setups
fdl probe --data-path /flodl/data        # override the shared-data path
fdl probe --libtorch-path /opt/libtorch  # override the libtorch directory
fdl probe --docker cuda         # NCCL is provided by a Docker image (compose service)

Exit code: 0 when every checked component is green, 1 when any issue was surfaced. The green path is silent enough to use as a post-deploy smoke test.

What it checks:

Output splits results into warnings (informational; do not block) and errors (block dispatch). fdl probe is a manual readiness gate — run it yourself before a cluster run; it is not invoked implicitly by fan-out. (The automatic pre-flight step fdl @cluster <cmd> performs is the per-host binary build, skippable with --no-prebuild.)

fdl status

Live status of a running cluster: lifecycle phase (waiting / staging / forming / training / done / failed), who has joined with what hardware, the join-window countdowns while it is still open, and — on start: manual / hybrid runs — the start switch’s state (a staging roster renders “roster startable, fire with fdl start). The controller serves the state as state.json over plain HTTP on its training port, so no extra port or config is involved — and curl works where fdl isn’t installed.

fdl @cluster status              # controller from the overlay's cluster.yml
fdl status --addr host[:port]    # explicit controller (default port 1337)
fdl @cluster status --json       # raw state.json for scripts
curl http://<controller>:1337/state.json   # same truth, no fdl

Address resolution: --addr wins; otherwise the active env’s cluster.controller (with a loopback retry, so all-tunneled runs are found when running on the controller box); otherwise the convention default 127.0.0.1:1337 (single-host auto-promoted runs), noted on stderr.

Exit code: 0 when the state was fetched and printed, 1 when no endpoint answered. The endpoint lives exactly as long as the launcher process — connection-refused after a run ends is the expected “no run listening” signal, not a fault.

fdl start

Fire the operator start switch of a staging cluster run. A join window opened with controller.join.start: manual (or hybrid) holds the roster once quorum is met instead of forming on the clock — the scavenged-credit shape: launch instances until the money runs out, watch fdl status, start when the roster looks full enough.

fdl @cluster start               # controller from the overlay's cluster.yml
fdl start --addr host[:port]     # explicit controller
fdl start --token <hex>          # non-loopback fire: the run's join token

Trust mirrors join admission: fired from the controller box (or through the sshd tunnel) the loopback peer address is the credential and no token is needed; from anywhere else --token must match the run’s controller.join.token. Address resolution is the same as fdl status.

Refusals name their reason and are never queued — auto mode (no switch to fire), quorum not met (with counts), window already closed, bad token. Exit code: 0 when the start was armed (the world forms at the next poll — watch fdl status), 1 otherwise.

fdl join

Join a cluster run’s window as a self-deployed worker: the worker-side walk-in for discovery windows (controller.join.discovery: true), where the window alone defines the world and worker addresses need not exist in any roster. It dials the controller, offers the box’s GPUs, and runs your training binary in agent role — the binary joins, then spawns and supervises this host’s relay and rank children itself; training code downstream is byte-identical to the fan-out path.

# Direct dial (trusted segment), authenticated by the run token:
fdl join 10.0.0.1:1337 --token <hex> --bin target/release/train -- --model resnet

# Through a guardrailed sshd on the controller box (the controller
# binds loopback under `tunnel_only`; reachability = authentication):
fdl join --ssh [email protected] --identity ~/.ssh/join_key \
         --bin target/release/train -- --model resnet

Every flag defaults from a top-level join: block in fdl.yml (see fdl.yml.example), so a golden image boots into bare fdl join.

Exit code: the agent’s own — 0 when this host’s ranks all finished cleanly. Under --persist the command only returns on setup errors. Full protocol walkthrough, trust model, and the join-sshd guardrail recipe: DDP reference.

fdl nccl

Build NVIDIA’s libnccl from source. Required for heterogeneous-rig clusters when the bundled NCCL versions across hosts don’t match (NCCL refuses handshake across major.minor skew). The build runs in a dedicated Docker context (Dockerfile.nccl.source) so no host-level NCCL/CUDA toolchain is required.

fdl nccl build                              # auto-detect target tag + local GPU archs
fdl nccl build --tag v2.27.5                # explicit NCCL git tag
fdl nccl build --archs "6.1;12.0"           # explicit archs (heterogeneous rig)
fdl nccl build --jobs 8                     # parallel compilation jobs (default 6)
fdl nccl build --dry-run                    # print build plan, do nothing

Auto-detection:

Output path: libtorch/nccl/builds/v<version>-<archs>/lib/libnccl.so.2.

Wire it into a worker via the env: LD_PRELOAD: block in fdl.cluster.yml:

workers:
  - host: node-b
    arch: builds/sm61-sm120
    env:
      LD_PRELOAD: /srv/flodl/libtorch/nccl/builds/v2.27.5-sm61/lib/libnccl.so.2

Build time: 5-15 minutes depending on CPU cores and arch count.

fdl --gpus

Scope GPU visibility for a single command. Two forms:

fdl --gpus all <cmd>            # use every visible CUDA device
fdl --gpus 0,1 <cmd>            # explicit physical indices

Behavior depends on the command kind:

Cluster context interaction: on fdl @cluster <cmd>, --gpus overrides per-worker local_devices: for the local controller host; remote workers continue to use their cluster.yml-declared devices. Loud-errors on duplicate, missing value, or invalid spec.

fdl --gpus 0 test                # CPU-style: scope tests to GPU 0
fdl --gpus 0,1 ddp-bench --mode nccl-cadence   # synthesize 2-rank single-host cluster
fdl --gpus all @cluster ddp-bench --mode nccl-cadence   # override local host devices in cluster mode

fdl @cluster <cmd> - multi-host fan-out

fdl @cluster <cmd> selects the cluster env overlay via the @ sigil. When fdl.cluster.yml exists alongside fdl.yml, it deep-merges as an overlay and triggers SSH fan-out for any command marked cluster: true.

Three equivalent forms:

fdl @cluster <cmd>           # sigil form
fdl --env cluster <cmd>      # explicit flag
FDL_ENV=cluster fdl <cmd>    # env-var form

Pre-flight per-host build: before fan-out, fdl @cluster <cmd> auto-builds the target binary locally for every remote host. Per-host CARGO_TARGET_DIR=target/cluster/<host>/<arch>/, libtorch resolved from each host’s arch: declaration, CUDA feature derived from the host’s GPU arch metadata. (Keying the target dir on arch as well as host means changing a host’s arch: rebuilds cleanly instead of reusing a binary linked against the old libtorch.) Builds run in parallel per host; first failure aborts fan-out. Remote dispatch invokes the prebuilt binary directly - no cargo, no rustc on remote.

Pass --no-prebuild to skip the pre-flight phase (when binaries are known fresh, or when iterating on a build-only issue).

Heterogeneous-rig flow (a Blackwell host + a Pascal VM, say):

  1. fdl @cluster probe - confirm GPU + libtorch + NCCL match per host.
  2. fdl nccl build on the host with the older NCCL - produces a matching libnccl.so.2 to wire via env: LD_PRELOAD:.
  3. fdl @cluster <cmd> - fan out.

See DDP Reference: Multi-host clusters for the fdl.cluster.yml schema and conventions.