HPC & Cluster Deployment
This page is the end-to-end path for running Novomodelo on a multi-node cluster: how to get an MPI-enabled binary, how to build the MPICH your cluster loads at runtime, how to launch under SLURM, and a worked AWS ParallelCluster reference architecture.
If you only run on a single machine you do not need any of this — use
--threads (see Running Studies)
and skip to Next Steps.
When you need MPI
Section titled “When you need MPI”Novomodelo parallelises on two levels, and they compose:
- MPI across ranks — distributes the forward-pass scenario batch and the backward-pass trial points across processes, typically one rank per node. This is the only way to scale past a single machine.
- rayon threads within a rank — solves LP subproblems concurrently inside
each rank. Controlled by
--threads.
The parallel model — how work is partitioned across ranks and threads, which
collectives tie the ranks together, the thread rules, and memory and case
access — is summarised in
Performance Accelerators — Parallel Execution Model;
the scheduler and communication internals live in the
novomodelo-sddp README.
This page is the operational counterpart: how to build and launch it.
A single novomodelo run uses the local (single-process) backend. Launching under an
MPI launcher (srun, mpiexec, or mpirun) switches Novomodelo to the MPI backend.
By default --comm-backend auto detects the launcher, so no flag is needed —
see Communication Backend.
Getting an MPI-enabled binary
Section titled “Getting an MPI-enabled binary”Option 1 — Pre-built novomodelo-mpi archive (recommended)
Section titled “Option 1 — Pre-built novomodelo-mpi archive (recommended)”Each Novomodelo release attaches a separate MPI archive alongside the standard binaries, named:
novomodelo-mpi-<version>-<target>.tar.gzDownload it from the GitHub Releases page.
The archive contains the binary (named novomodelo-mpi to distinguish it from the
single-process novomodelo in the same release), the license/notice files, and a
README.txt. The README gives its SLURM launch lines with srun --mpi=pmi2;
its section on stopping before a SLURM time limit and its troubleshooting notes
use srun --mpi=pmix. With
an MPICH built for PMIx, as on this page, launch with srun --mpi=pmix (see
Running under SLURM).
| Platform | Target triple |
|---|---|
| Linux (x86-64) | x86_64-unknown-linux-gnu |
| Linux (ARM64) | aarch64-unknown-linux-gnu |
| macOS (Apple Silicon) | aarch64-apple-darwin |
Windows and Intel macOS have no pre-built MPI archive — build from source (Option 2) on those platforms.
The x86-64 Linux novomodelo-mpi uses AVX2 and FMA instructions and needs a
processor that supports them.
The Linux binaries are dynamically linked against libmpi.so.12 (the MPICH 4.x
ABI). They do not bundle an MPI runtime — you supply one at run time (see
Building MPICH from source). Any MPICH-ABI-compatible
runtime works:
- MPICH 4.0+ (recommended)
- Intel MPI 2021+
- MVAPICH2 3.0+
Option 2 — Build from source with the mpi feature
Section titled “Option 2 — Build from source with the mpi feature”For a platform without a pre-built archive, or when you want to link against a
specific MPI, build the CLI with the mpi feature:
# MPICH (mpicc) must be on PATH and discoverable via pkg-config first —# see "Building MPICH from source" below.cargo build --release --features mpi -p novomodelo-cliThe binary is written to target/release/novomodelo. The build reads the MPI
headers and link flags through pkg-config, so ensure PKG_CONFIG_PATH
includes your MPICH lib/pkgconfig directory.
On x86-64 Linux the source tree’s .cargo/config.toml enables AVX2 and FMA for
the build, so the binary needs a processor that supports them, as the pre-built
archive does, unless a RUSTFLAGS environment variable replaces that setting.
Building MPICH from source
Section titled “Building MPICH from source”On an HPC cluster you almost always build MPICH yourself so it links against the cluster’s high-performance fabric and process manager, rather than a generic package. The recipe below is the one the AWS ParallelCluster reference architecture uses; it is pinned to the same MPICH version the release binaries are compiled against.
MPICH_VERSION=4.2.3
wget "https://www.mpich.org/static/downloads/${MPICH_VERSION}/mpich-${MPICH_VERSION}.tar.gz"tar xzf "mpich-${MPICH_VERSION}.tar.gz"cd "mpich-${MPICH_VERSION}"
./configure \ --prefix=/opt/mpich \ --with-device=ch4:ofi \ --with-libfabric=/opt/amazon/efa \ --with-pmi=pmix \ --with-pmix=/opt/pmix \ --with-slurm=/opt/slurm \ --with-pm=no
make -j"$(nproc)"make installBefore you use the build, confirm that libmpi links the cluster’s PMIx and
not the built-in client:
ldd /opt/mpich/lib/libmpi.so.12 | grep libpmix# libpmix.so.2 => /opt/pmix/lib/libpmix.so.2 (0x...)If this prints nothing, or the library resolves outside the PMIx that SLURM
uses, reconfigure with the correct --with-pmix path and rebuild. PMIx 4.x and
5.x both ship libpmix.so.2, so the thing to check is the resolved path,
not the soname.
Then make it discoverable for both compiling (pkg-config) and running
(LD_LIBRARY_PATH):
export PATH=/opt/mpich/bin:$PATHexport LD_LIBRARY_PATH=/opt/mpich/lib:${LD_LIBRARY_PATH:-}export PKG_CONFIG_PATH=/opt/mpich/lib/pkgconfig:${PKG_CONFIG_PATH:-}What each flag does:
| Flag | Purpose |
|---|---|
--with-device=ch4:ofi | Use the modern CH4 device over an OpenFabrics Interfaces (libfabric) provider — the fabric path. |
--with-libfabric=/opt/amazon/efa | Link against the EFA installer’s libfabric so the ofi netmod can select the EFA provider. |
--with-pmi=pmix | Use PMIx as the only process-management interface, so srun --mpi=pmix launches and wires up the ranks. No PMI-1/PMI-2 client is built. |
--with-pmix=/opt/pmix | Link the cluster’s PMIx library, the same one SLURM’s pmix plugin uses. If you omit it, MPICH silently uses its built-in client (see the caution above). |
--with-slurm=/opt/slurm | Integrate with the cluster’s SLURM installation. |
--with-pm=no | Build no internal process manager (no Hydra/mpiexec) — SLURM launches the ranks. |
The paths above (/opt/amazon/efa, /opt/pmix, /opt/slurm) are the
conventional locations for the EFA installer, PMIx, and SLURM on AWS
ParallelCluster; adjust them to your environment.
If you cannot write to /opt, install into a prefix on a filesystem every node
can see, for example --prefix="$HOME/mpich-4.2.3". Then use that prefix in the
PATH/LD_LIBRARY_PATH exports and in the batch script. Replacing the MPICH
runtime does not require rebuilding novomodelo-mpi, because the binary only
depends on the stable libmpi.so.12 ABI.
Optional build flags:
- Trims. Novomodelo uses no Fortran, C++, or MPI-IO bindings, so
--disable-fortran --disable-cxx --disable-romio --disable-docshorten the build without affecting Novomodelo. - Faster runtime.
--enable-fast=O3,ndebugbuilds an optimised MPICH without internal assertions.
Running under SLURM
Section titled “Running under SLURM”On a SLURM cluster, srun is the launcher: it places the ranks, binds them to
cores, and (via PMIx) initialises MPI. Because --comm-backend defaults to
auto, Novomodelo detects the srun launch and selects the MPI backend
automatically.
The commands below use three placeholders. <partition> is the SLURM partition,
<shared-path> is a directory every node can read that holds novomodelo-mpi and the
batch script, and <case-dir> is the case directory on a shared filesystem, at
the same path on every node because every rank reads it.
Hybrid MPI + threads
Section titled “Hybrid MPI + threads”Novomodelo uses MPI for inter-node communication and rayon threads for intra-node LP solves. The launch below runs one rank per node, with the thread pool sized as set out in Sizing and scaling:
srun --mpi=pmix --nodes=4 --ntasks-per-node=1 --cpus-per-task=96 --exclusive \ ./novomodelo-mpi run <case-dir> --threads 96In this command --threads 96 matches --cpus-per-task=96, so each rank’s
rayon pool has one thread per allocated core. --exclusive gives each rank a
whole node with no co-tenants competing for cores or memory bandwidth. The 4
nodes and the 96 cores are illustrative: set them to the allocation your job
needs.
Batch script
Section titled “Batch script”A production job wraps that launch in an sbatch script. Replace the
placeholders and set the node count, cores per task and time limit for your
workload:
#!/bin/bash#SBATCH --job-name=novomodelo-mpi#SBATCH --nodes=4#SBATCH --ntasks-per-node=1#SBATCH --cpus-per-task=96#SBATCH --exclusive#SBATCH --output=novomodelo_%j.out#SBATCH --error=novomodelo_%j.err#SBATCH --partition=<partition>#SBATCH --time=0-08:00:00
MPICH_PREFIX=/opt/mpichNOVOMODELO_MPI="<shared-path>/novomodelo-mpi"
export PATH="$MPICH_PREFIX/bin:$PATH"export LD_LIBRARY_PATH="$MPICH_PREFIX/lib:${LD_LIBRARY_PATH:-}"
srun --mpi=pmix "$NOVOMODELO_MPI" run "$1" --threads "$SLURM_CPUS_PER_TASK"Refer to the binary by an absolute path. sbatch runs the script in the
directory you submitted from, so ./novomodelo-mpi works only if you happen to run
sbatch from the directory that holds the binary.
Save the script as run-novomodelo.sh in <shared-path>, next to novomodelo-mpi. Submit
it with the case directory as its argument:
sbatch <shared-path>/run-novomodelo.sh <case-dir>Key SLURM flags
Section titled “Key SLURM flags”| Flag | Purpose |
|---|---|
--mpi=pmix | PMIx process startup (recommended; see the tip below) |
--mpi=pmi2 | PMI-2 startup. It needs an MPICH built with PMI-2 support; the --with-pmi=pmix build above has none |
--nodes=N | Number of nodes |
--ntasks-per-node=N | MPI ranks per node (1 for the one-rank-per-node mapping) |
--cpus-per-task=T | Cores per rank — pass the same value to --threads |
--exclusive | Give each rank a whole node |
--cpu-bind=cores | Pin each rank’s threads to specific cores |
--mem-bind=local | Allocate memory from the NUMA node closest to the bound cores |
Sizing and scaling
Section titled “Sizing and scaling”No measured scaling data is published, so these rules relate ranks, threads, forward passes and memory to one another instead of giving node or core counts; Performance Accelerators — Parallel Execution Model holds the facts behind them.
- Run one rank per node, with
--threadsequal to the cores allocated to the rank (--cpus-per-task). Each LP solve runs on one thread, so a rank solves at most--threadsLPs at a time. - Run every rank with the same
--threads. Under a sampled forward pass, the run stops with an error when ranks differ in--threads. - Under the
sampledmethod, settraining.selection.forward_passesandsimulation.selection.num_scenariosto at least the rank count times--threads, and to a multiple of it for equal shares. On more than one rank with--threadsabove1, aforward_passesthat is not a multiple of the rank count stops one rank with anindex out of boundserror while the other ranks wait, so the job runs until its wall time. Trajectories and scenarios are split over the ranks and then over each rank’s workers, and a worker left without work idles. - Size each node’s memory for one rank’s full copy of the case, the opening tree, the cut pool and the stage templates. Each rank holds its own copy, so each extra rank on a node adds another, and the cut pool scales with the iteration budget and the forward passes.
- Before adding ranks or threads, read the printed
Time splitand the knob table in Performance Accelerators — Diagnosing Performance. The lower-bound evaluation runs serially on rank 0, and under a sampled forward pass every rank waits on the per-stage cut exchange, so neither shrinks with more ranks or threads. - Place the case directory as the
<case-dir>definition under Running under SLURM requires.
Resuming a job
Section titled “Resuming a job”A job writes its checkpoint when training ends and, with periodic checkpoints on,
on each scheduled iteration. A job killed at its --time
loses the iterations after its last checkpoint, and the next job resumes from that checkpoint.
An interrupted write that replaces a checkpoint leaves a complete one, old or new; an interrupted first write can leave none
(see Resume).
To fit a long training into wall-time-limited jobs, split it into slices, one
batch job each, and submit the next slice after the previous one ends. Every
slice after the first sets policy.mode to "resume" and raises
iteration_limit to the new total, which counts the iterations of all earlier
slices (Train in Slices gives the
rule sets and how to size a slice).
The job’s --time covers the case load, the load of the previous checkpoint, the
slice’s iterations and the checkpoint write, plus the simulation and its result
export in a slice that enables simulation (simulation.enabled set to true; the
default is off). The checkpoint is written before simulation starts, so a job
killed during simulation keeps that slice’s checkpoint.
After editing <case-dir>/config.json as above, submit each slice
on the same <case-dir>. The batch script passes no --output, so every slice uses <case-dir>/output on the shared
filesystem, and every rank loads the previous checkpoint from it:
sbatch <shared-path>/run-novomodelo.sh <case-dir>AWS ParallelCluster reference architecture
Section titled “AWS ParallelCluster reference architecture”Novomodelo’s production deployment runs on AWS ParallelCluster: a lightweight, always-on head node for submission plus one or more compute queues whose nodes are provisioned on demand and torn down when idle. Each queue is an independent SLURM partition that scales on its own.
The moving parts:
- Custom machine image (shared by head and compute nodes). MPICH is built
from source into the image at
/opt/mpich, using thech4:ofi+ EFA recipe above. Baking it into the image means every dynamically-launched compute node already has the runtime — no per-job install step. - The binary. The pre-built
novomodelo-mpiarchive is downloaded from the release page to<shared-path>. It is the same binary produced by the release CI, and the custom image’s system libraries are kept compatible with that build. - Interconnect (EFA). Because MPICH is built
--with-libfabric=/opt/amazon/efaand--with-device=ch4:ofi, the OFI netmod selects the EFA provider on EFA-enabled instances automatically. For fabric-level diagnostics or provider overrides, consult the AWS EFA and libfabric documentation. - Instances. The compute queue in the example uses
c7a.48xlargenodes with the one-rank-per-node ×--threads=$SLURM_CPUS_PER_TASKmapping; smaller compute-optimised instances work the same way. - Storage (FSx). The case directory, checkpoints, and exported results live
on the shared FSx filesystem. Novomodelo’s per-solve hot path holds its working set
in memory (its allocation discipline is documented in the
novomodelo-sddpREADME); it touches the shared filesystem heavily only at checkpointing and result export, so FSx throughput is sized for those phases rather than the solve loop.
Submission is the same sbatch <shared-path>/run-novomodelo.sh <case-dir> shown above:
the head node queues the job, ParallelCluster powers up the compute nodes in the
target partition, srun --mpi=pmix launches one rank per node, and the nodes
scale back down when the queue drains.
Troubleshooting
Section titled “Troubleshooting”error while loading shared libraries: libmpi.so.12 — the MPI runtime is
not on the loader path. Export LD_LIBRARY_PATH=/opt/mpich/lib:$LD_LIBRARY_PATH
(or module load mpich) before launching, as the batch script does.
comm: local when you expected comm: mpi — either the binary was built
without the mpi feature (check novomodelo version), or the process was not
launched under a recognised launcher. Force the backend with
--comm-backend mpi; it fails with a clear message on a non-MPI binary rather
than silently running single-process.
Every rank reports Layout: 1 rank, and the run repeats once per node.
Symptoms: the Execution banner appears once per rank, not once in total, and
each copy reads Layout: 1 rank on <host>. The simulation reports
... across 1 ranks, and the separate copies collide on the shared output
directory with write errors. MPI_Init fell back to a singleton: every
process launched by srun became its own one-rank MPI_COMM_WORLD.
novomodelo-mpi version still reports comm: mpi and the banner still shows
Backend: MPI, because the binary is fine. The fault is in the MPICH runtime:
it was built without --with-pmix, or against a different PMIx than SLURM’s
plugin, so the PMIx handshake failed silently.
To fix it:
- Confirm the cause with
ldd <mpich-prefix>/lib/libmpi.so.12 | grep libpmix. Seeing no line, or a path outside the cluster’s PMIx, confirms it. - Rebuild MPICH with
--with-pmixpointing at the PMIx that SLURM uses (see Building MPICH from source). - Point the batch script’s exports at the new build.
novomodelo-mpiitself needs no rebuild.
Switching to --mpi=pmi2 does not fix this.
srun --mpi=pmi2 fails with pmijobid missing in fullinit command. An
MPICH built --with-pmi=pmix (the recommended build) contains no PMI-2 client,
so --mpi=pmi2 is not an alternative launch path for it. Use --mpi=pmix. The
same error is also caused by an incompatibility between SLURM 24.05’s PMI2
plugin and MPICH’s PMI2 wire protocol. If only PMI2 is available and you built
Hydra (i.e. not --with-pm=no), mpiexec bypasses PMI2 entirely.
Multi-node job hangs at MPI_Init on MPICH 4.3.x — the 4.3.0 PMI2 client
regression described above. Downgrade to MPICH 4.2.x.
Verifying your setup
Section titled “Verifying your setup”For one novomodelo-mpi binary, with every node running the same system image on one
processor model (the same libm, libstdc++ and MPI library), the numerical content
of the output files is bit-identical at any rank count and at any thread count
(the same --threads on every rank); timestamps, wall-clock timing columns and
the recorded execution topology differ. The guarantee
holds for a stopping-rule set without time_limit: a wall-clock or signal stop depends on timing
(time_limit). Validate on one rank with
novomodelo-mpi itself, because the single-process novomodelo is a different build
whose results are not promised equal. The full statement is in
Determinism & Provenance.
Check your setup in three steps. Each step catches a different failure.
-
The binary is MPI-enabled. This only checks the build, not the runtime wiring:
Terminal window <shared-path>/novomodelo-mpi version # → comm: mpi -
MPICH links the cluster’s PMIx. Run the
ldd … | grep libpmixcheck from Building MPICH from source. -
The ranks form one job. Run a two-node smoke test, which exercises both PMIx wire-up and the inter-node fabric:
Terminal window srun --mpi=pmix --nodes=2 --ntasks-per-node=1 \<shared-path>/novomodelo-mpi run <case-dir> --threads 2Only rank 0 prints the
Executionbanner, so it must appear once. ItsBackendandLayoutlines on a multi-node job have this format, followed by one host line per node ({rank_word}isrankorranks;{range}lists the host’s ranks, such as0–1):Backend: MPI ({library_version}, {standard_version})Layout: {world_size} {rank_word} across {num_hosts} nodes{hostname}: ranks {range} ({count} {rank_count_word})If the banner appears twice, each copy reading
Layout: 1 rank on <host>, the ranks fell back to singletons. See Troubleshooting.
See Also
Section titled “See Also”- Performance Accelerators — the parallel-execution model MPI drives
- Running Studies — the
--comm-backendflag and single-node threading - CLI Reference — complete flag and subcommand reference
- Determinism & Provenance — the per-binary reproducibility guarantee
- Installation — standard (single-process) install paths