Project and libtorch Commands

Interactive wizard that walks you through everything:

  1. Detects your system - CPU, RAM, Docker, Rust, GPUs.
  2. Downloads libtorch - auto-picks the right variant for your GPU(s).
  3. Configures your build - Docker or native, builds images if needed.
fdl setup                      # interactive (asks questions)
fdl setup --non-interactive    # auto-detect everything, no prompts
fdl setup -y                   # alias for --non-interactive
fdl setup --force              # re-download even if libtorch exists

The wizard handles tricky scenarios automatically:

fdl libtorch

Manage libtorch installations. Variants live under libtorch/ in your project (or $FLODL_HOME/libtorch/ when standalone), each with a metadata .arch file. An .active pointer selects the current one.

fdl libtorch download

Download a pre-built libtorch from PyTorch’s official mirrors.

fdl libtorch download              # auto-detect GPU, pick best variant
fdl libtorch download --cpu        # force CPU-only (~200MB)
fdl libtorch download --cuda 12.8  # CUDA 12.8 / cu128 (~2GB)
fdl libtorch download --cuda 12.6  # CUDA 12.6 / cu126 (~2GB)
fdl libtorch download --rocm 7.0   # AMD ROCm 7.0 (~5GB, Linux only)
fdl libtorch download --rocm 7.1   # AMD ROCm 7.1 (same GPUs, newer runtime)
fdl libtorch download --path ~/lib # install to a custom directory
fdl libtorch download --no-activate # install but do not switch `.active`
fdl libtorch download --dry-run    # show what would happen

--cuda only accepts 12.6 or 12.8 and --rocm only 7.0 or 7.1 (the published pre-built versions). Auto-completion offers each.

Variant coverage:

Variant Architectures GPUs
CPU - Any (no GPU acceleration)
cu126 sm_50 to sm_90 Maxwell through Ada Lovelace
cu128 sm_70 to sm_120 Volta through Blackwell
rocm70 gfx908 to gfx1201 CDNA 1-4 (MI100/MI200/MI300/MI350), RDNA 2-4 incl. Strix APUs
rocm71 same as rocm70 same as rocm70, newer HIP runtime

A libtorch build serves exactly one vendor, so there is no variant covering both an NVIDIA and an AMD card in one process. On a box holding both, auto-detect picks the NVIDIA cards and prints the --rocm line to use instead.

Which ROCm version. The two ROCm variants reach exactly the same cards, so the choice is about the runtime. flodl puts the host’s own ROCm ahead of the bundled one on LD_LIBRARY_PATH (a bundled HIP runtime that disagrees with the host’s kernel driver segfaults at the first GPU op), and within a major version that ABI only grows: a bundle older than the host loads, a newer one can fail on a symbol the host does not have. So auto-detect picks 7.0, which serves every ROCm 7.x host, and --rocm 7.1 is there for a host that wants the exact match.

The ROCm archive ships pre-built rocBLAS kernels for a fixed gfx list:

gfx908 gfx90a gfx942 gfx950 gfx1030 gfx1100 gfx1101 gfx1102 gfx1150
gfx1151 gfx1200 gfx1201

gfx900 and gfx906 (Vega 10 / Vega 20, i.e. MI50 and Radeon VII) are not on it: the archive carries MIOpen tuning databases for them but no rocBLAS kernels, so they are treated as uncovered rather than installed and left to fail at the first matmul.

A card outside it has no kernels, so auto-detect stays on CPU and names the covered targets rather than installing something that cannot run. ROCm libtorch is published for Linux only; --rocm on Windows or macOS is refused with that reason.

When an AMD box behaves in a way fdl probe / fdl diagnose cannot explain, rocm-capture.sh at the repo root snapshots the whole stack in one pass (KFD topology, ROCm userspace, device permissions, both sides of the container boundary) — attach its output to a bug report.

If your NVIDIA GPUs span both CUDA ranges (e.g. GTX 1060 + RTX 5060 Ti), no single pre-built variant covers both. Use fdl libtorch build.

fdl libtorch build

Compile libtorch from PyTorch source for your exact GPU combination. Takes 2-6 hours depending on CPU cores.

This is a CUDA-only path: --archs takes compute capabilities and the toolchain expects nvcc. AMD is served by the pre-built --rocm 7.0 download, which already covers every gfx target upstream builds kernels for; there is no source-build equivalent.

Two build methods are available:

When both are available, the CLI asks which you prefer. Use --docker or --native to skip the prompt.

fdl libtorch build                         # auto-detect GPUs and backend
fdl libtorch build --native                # force native build
fdl libtorch build --docker                # force Docker build
fdl libtorch build --archs "6.1;12.0"      # explicit architectures
fdl libtorch build --jobs 8                # parallel compilation jobs (default: 6)
fdl libtorch build --dry-run               # show plan without building

Output lands in libtorch/builds/<arch-signature>/ (e.g. libtorch/builds/sm61-sm120/).

Native build requirements:

Tool Purpose Install
nvcc CUDA compiler CUDA Toolkit
cmake Build system apt install cmake / brew install cmake
python3 PyTorch build scripts Usually pre-installed
git Clone PyTorch source apt install git
gcc/g++ C++ compilation apt install gcc g++

Python packages (pyyaml, jinja2, etc.) install automatically via pip. The PyTorch source is cached at libtorch/.build-cache/pytorch/, so re-running after a failure skips the clone.

fdl libtorch list / info / activate / remove

fdl libtorch list            # human-readable
fdl libtorch list --json     # machine-readable
fdl libtorch info            # show active variant details
fdl libtorch activate <name> # switch the active variant
fdl libtorch remove <name>   # delete a variant (clears .active if it was active)

activate and remove take a variant name as shown by fdl libtorch list (e.g. precompiled/cu128, builds/sm61-sm120). Passing no name prints the list and exits.

Example info output:

Active:   builds/sm61-sm120
Version:  2.10.0
CUDA:     12.8
Archs:    6.1 12.0
Source:   compiled

The same fields describe an AMD variant, with CUDA: none (the field is the CUDA toolkit version, which a ROCm build has none of) and gfx targets in Archs:

Active:   precompiled/rocm70
Version:  2.10.0
CUDA:     none
Archs:    gfx906 gfx908 gfx90a gfx942 gfx1030 gfx1100 gfx1101 gfx1102 gfx1200 gfx1201
Source:   precompiled

Using fdl as a standalone libtorch manager (tch-rs / PyTorch C++)

The libtorch-management and diagnostics commands are independent of flodl and fill a gap PyTorch itself never filled: a proper installer. fdl works as a drop-in libtorch manager for:

Standalone (no project directory), everything installs under $FLODL_HOME (default ~/.flodl/). Pick any location you prefer and export it before the first command:

export FLODL_HOME=~/.libtorch-variants

Example A: PyTorch C++ (LibTorch via CMake) on an RTX 50-series GPU.

This is the canonical C++ API workflow from pytorch.org/cppdocs, with fdl replacing the manual URL-and-unzip dance:

# 1. Inspect hardware and download the matching libtorch.
fdl diagnose                          # confirm GPU arch (sm_120 in this case)
fdl libtorch download --cuda 12.8     # ~2GB, unpacks to $FLODL_HOME/libtorch/precompiled/cu128

# 2. Point CMake at it.
export LIBTORCH=$FLODL_HOME/libtorch/precompiled/cu128

Minimal CMakeLists.txt:

cmake_minimum_required(VERSION 3.18 FATAL_ERROR)
project(my_model)

find_package(Torch REQUIRED)
add_executable(my_model main.cpp)
target_link_libraries(my_model "${TORCH_LIBRARIES}")
set_property(TARGET my_model PROPERTY CXX_STANDARD 17)

Build and run:

mkdir build && cd build
cmake -DCMAKE_PREFIX_PATH=$LIBTORCH ..
cmake --build . --parallel

# Runtime: expose libtorch's shared libs.
export LD_LIBRARY_PATH=$LIBTORCH/lib:$LD_LIBRARY_PATH
./my_model

To switch CUDA versions (e.g. back to 12.6 for legacy code), install the other variant with fdl libtorch download --cuda 12.6, flip it with fdl libtorch activate precompiled/cu126, re-export LIBTORCH, and re-run CMake. No reinstall, no URL hunting.

Example B: Rust via tch-rs on the same hardware.

# Same download + LIBTORCH export as Example A.
export LIBTORCH=$FLODL_HOME/libtorch/precompiled/cu128
export LD_LIBRARY_PATH=$LIBTORCH/lib:$LD_LIBRARY_PATH
cargo add tch
cargo build

Juggling variants across projects. Install as many as you need side by side, then flip the active pointer; LIBTORCH follows .active when you source it from the fdl libtorch info output:

fdl libtorch download --cpu           # ~200MB, for laptops / CI
fdl libtorch download --cuda 12.6     # legacy CUDA projects
fdl libtorch download --cuda 12.8     # latest

fdl libtorch activate precompiled/cu126   # work on legacy code
fdl libtorch activate precompiled/cu128   # work on RTX 50-series code
fdl libtorch info                         # confirm what's active

Mixed GPUs (no pre-built variant covers you). If fdl diagnose reports architectures that span both pre-built ranges, build from source and fdl will pick up the compiled variant automatically:

fdl libtorch build --archs "6.1;12.0"     # Pascal + Blackwell
fdl libtorch list
#   builds/sm61-sm120 (active)
#   precompiled/cu128
export LIBTORCH=$FLODL_HOME/libtorch/builds/sm61-sm120

CI gating example. Use diagnose --json to skip GPU jobs when no compatible device is present:

if fdl diagnose --json | jq -e '.cuda.devices | length > 0' > /dev/null; then
    cargo test --features cuda
else
    echo "no GPU detected, skipping CUDA tests"
fi

None of the above touches flodl itself - fdl is just the libtorch installer / activator / diagnostics tool in this mode.

fdl init

Scaffold a new floDl project. Three modes, mutually exclusive - pick via flag, or accept the interactive prompt when none is passed:

fdl init my-model            # default: Docker with host-mounted libtorch (prompts if interactive)
fdl init my-model --docker   # Docker with libtorch baked into the image
fdl init my-model --native   # no Docker; libtorch and cargo on the host

Add --with-hf to include the flodl-hf HuggingFace playground in the generated project:

fdl init my-model --with-hf            # Docker + flodl-hf side crate
fdl init my-model --native --with-hf   # Native + flodl-hf side crate

--with-hf skips the interactive “Include flodl-hf?” prompt when mode flags are present. In fully interactive mode (fdl init my-model with no flag), a prompt offers the same choice after the Docker / native selection. See fdl add below for adding flodl-hf to an existing project later.

In all three modes the scaffold generates:

Docker modes additionally generate:

Native mode skips all the Docker files - commands run on the host. Point $LIBTORCH / $LD_LIBRARY_PATH at a libtorch install (use ./fdl libtorch download --cpu, --cuda 12.8 or --rocm 7.0) and ./fdl build dispatches straight to cargo build.

The scaffold is fdl-native: there is no Makefile. Every task lives in fdl.yml and runs via ./fdl <cmd>. Libtorch environment variables (LIBTORCH_HOST_PATH, CUDA_VERSION, CUDA_TAG) are derived from libtorch/.active by flodl-cli before each dispatch - the logic that used to live in the scaffolded Makefile now lives in one place inside the binary.

fdl add

Add an ecosystem crate as a side playground inside an initialised flodl project. Today this means flodl-hf (alias hf); the command is designed to grow as more sibling crates land.

fdl add flodl-hf             # scaffold ./flodl-hf/
fdl add hf                   # short alias, same effect

The scaffold drops a standalone cargo crate under ./flodl-hf/ with its own Cargo.toml, a one-file AutoModel classifier (src/main.rs), a nested fdl.yml with runnable commands (classify, bert, roberta-sentiment, distilbert-sentiment, plus build / check / shell), and a README covering the three feature flavors (full / vision-only / offline) and the .bin-to-safetensors conversion workflow.

Key properties:

See the HuggingFace Integration tutorial for the full usage walkthrough of what the scaffold enables.

fdl diagnose

Hardware and compatibility report. Useful for debugging setup issues or verifying your GPU + libtorch combination works.

fdl diagnose             # human-readable report
fdl diagnose --json      # machine-readable for CI and tooling

Example output:

floDl Diagnostics
=================

System
  CPU:         Intel(R) Core(TM) i9-9900K CPU @ 3.60GHz (16 threads, 24GB RAM)
  OS:          Linux 6.6.87.2-microsoft-standard-WSL2 (WSL2)
  Docker:      29.3.1

CUDA
  Driver:      576.88
  Devices:     2
  [0] NVIDIA GeForce RTX 5060 Ti -- sm_120, 15GB VRAM
  [1] NVIDIA GeForce GTX 1060 6GB -- sm_61, 6GB VRAM

libtorch
  Active:      builds/sm61-sm120
  Version:     2.10.0
  CUDA:        12.8
  Archs:       6.1 12.0
  Source:      compiled
  Variants:    builds/sm61-sm120, precompiled/cpu

Compatibility
  GPU 0 (RTX 5060 Ti, sm_120):  OK
  GPU 1 (GTX 1060 6GB, sm_61):  OK

  All GPUs compatible with active libtorch.

The JSON output is useful for CI pipelines and automated tooling:

fdl diagnose --json | jq '.cuda.devices[] | .sm'

libtorch directory layout

The CLI manages libtorch installations under libtorch/ in your project root:

libtorch/
  .active                          # points to current variant (e.g. "builds/sm61-sm120")
  precompiled/
    cpu/                           # pre-built CPU variant
      lib/ include/ share/
      .arch                        # metadata: cuda=none, torch=2.10.0, ...
    cu126/                         # pre-built CUDA 12.6
      ...
    cu128/                         # pre-built CUDA 12.8
      ...
  builds/
    sm61-sm120/                    # source-built for specific GPUs
      lib/ include/ share/
      .arch                        # metadata: cuda=12.8, archs=6.1 12.0, source=compiled

The .arch file format:

cuda=12.8
torch=2.10.0
archs=6.1 12.0
source=compiled

Docker Compose and Make targets read .active to mount the right libtorch variant automatically. You never need to set LIBTORCH_PATH manually when using Docker.