Skip to content

HPC & Cluster Deployment

This page is the end-to-end path for running Novomodelo on a multi-node cluster: how to get an MPI-enabled binary, how to build the MPICH your cluster loads at runtime, how to launch under SLURM, and a worked AWS ParallelCluster reference architecture.

If you only run on a single machine you do not need any of this — use --threads (see Running Studies) and skip to Next Steps.


Novomodelo parallelises on two levels, and they compose:

  • MPI across ranks — distributes the forward-pass scenario batch and the backward-pass trial points across processes, typically one rank per node. This is the only way to scale past a single machine.
  • rayon threads within a rank — solves LP subproblems concurrently inside each rank. Controlled by --threads.

The parallel model — how work is partitioned across ranks and threads, which collectives tie the ranks together, the thread rules, and memory and case access — is summarised in Performance Accelerators — Parallel Execution Model; the scheduler and communication internals live in the novomodelo-sddp README. This page is the operational counterpart: how to build and launch it.

A single novomodelo run uses the local (single-process) backend. Launching under an MPI launcher (srun, mpiexec, or mpirun) switches Novomodelo to the MPI backend. By default --comm-backend auto detects the launcher, so no flag is needed — see Communication Backend.


Section titled “Option 1 — Pre-built novomodelo-mpi archive (recommended)”

Each Novomodelo release attaches a separate MPI archive alongside the standard binaries, named:

novomodelo-mpi-<version>-<target>.tar.gz

Download it from the GitHub Releases page. The archive contains the binary (named novomodelo-mpi to distinguish it from the single-process novomodelo in the same release), the license/notice files, and a README.txt. The README gives its SLURM launch lines with srun --mpi=pmi2; its section on stopping before a SLURM time limit and its troubleshooting notes use srun --mpi=pmix. With an MPICH built for PMIx, as on this page, launch with srun --mpi=pmix (see Running under SLURM).

PlatformTarget triple
Linux (x86-64)x86_64-unknown-linux-gnu
Linux (ARM64)aarch64-unknown-linux-gnu
macOS (Apple Silicon)aarch64-apple-darwin

Windows and Intel macOS have no pre-built MPI archive — build from source (Option 2) on those platforms.

The x86-64 Linux novomodelo-mpi uses AVX2 and FMA instructions and needs a processor that supports them.

The Linux binaries are dynamically linked against libmpi.so.12 (the MPICH 4.x ABI). They do not bundle an MPI runtime — you supply one at run time (see Building MPICH from source). Any MPICH-ABI-compatible runtime works:

  • MPICH 4.0+ (recommended)
  • Intel MPI 2021+
  • MVAPICH2 3.0+

Option 2 — Build from source with the mpi feature

Section titled “Option 2 — Build from source with the mpi feature”

For a platform without a pre-built archive, or when you want to link against a specific MPI, build the CLI with the mpi feature:

Terminal window
# MPICH (mpicc) must be on PATH and discoverable via pkg-config first —
# see "Building MPICH from source" below.
cargo build --release --features mpi -p novomodelo-cli

The binary is written to target/release/novomodelo. The build reads the MPI headers and link flags through pkg-config, so ensure PKG_CONFIG_PATH includes your MPICH lib/pkgconfig directory.

On x86-64 Linux the source tree’s .cargo/config.toml enables AVX2 and FMA for the build, so the binary needs a processor that supports them, as the pre-built archive does, unless a RUSTFLAGS environment variable replaces that setting.


On an HPC cluster you almost always build MPICH yourself so it links against the cluster’s high-performance fabric and process manager, rather than a generic package. The recipe below is the one the AWS ParallelCluster reference architecture uses; it is pinned to the same MPICH version the release binaries are compiled against.

Terminal window
MPICH_VERSION=4.2.3
wget "https://www.mpich.org/static/downloads/${MPICH_VERSION}/mpich-${MPICH_VERSION}.tar.gz"
tar xzf "mpich-${MPICH_VERSION}.tar.gz"
cd "mpich-${MPICH_VERSION}"
./configure \
--prefix=/opt/mpich \
--with-device=ch4:ofi \
--with-libfabric=/opt/amazon/efa \
--with-pmi=pmix \
--with-pmix=/opt/pmix \
--with-slurm=/opt/slurm \
--with-pm=no
make -j"$(nproc)"
make install

Before you use the build, confirm that libmpi links the cluster’s PMIx and not the built-in client:

Terminal window
ldd /opt/mpich/lib/libmpi.so.12 | grep libpmix
# libpmix.so.2 => /opt/pmix/lib/libpmix.so.2 (0x...)

If this prints nothing, or the library resolves outside the PMIx that SLURM uses, reconfigure with the correct --with-pmix path and rebuild. PMIx 4.x and 5.x both ship libpmix.so.2, so the thing to check is the resolved path, not the soname.

Then make it discoverable for both compiling (pkg-config) and running (LD_LIBRARY_PATH):

Terminal window
export PATH=/opt/mpich/bin:$PATH
export LD_LIBRARY_PATH=/opt/mpich/lib:${LD_LIBRARY_PATH:-}
export PKG_CONFIG_PATH=/opt/mpich/lib/pkgconfig:${PKG_CONFIG_PATH:-}

What each flag does:

FlagPurpose
--with-device=ch4:ofiUse the modern CH4 device over an OpenFabrics Interfaces (libfabric) provider — the fabric path.
--with-libfabric=/opt/amazon/efaLink against the EFA installer’s libfabric so the ofi netmod can select the EFA provider.
--with-pmi=pmixUse PMIx as the only process-management interface, so srun --mpi=pmix launches and wires up the ranks. No PMI-1/PMI-2 client is built.
--with-pmix=/opt/pmixLink the cluster’s PMIx library, the same one SLURM’s pmix plugin uses. If you omit it, MPICH silently uses its built-in client (see the caution above).
--with-slurm=/opt/slurmIntegrate with the cluster’s SLURM installation.
--with-pm=noBuild no internal process manager (no Hydra/mpiexec) — SLURM launches the ranks.

The paths above (/opt/amazon/efa, /opt/pmix, /opt/slurm) are the conventional locations for the EFA installer, PMIx, and SLURM on AWS ParallelCluster; adjust them to your environment.

If you cannot write to /opt, install into a prefix on a filesystem every node can see, for example --prefix="$HOME/mpich-4.2.3". Then use that prefix in the PATH/LD_LIBRARY_PATH exports and in the batch script. Replacing the MPICH runtime does not require rebuilding novomodelo-mpi, because the binary only depends on the stable libmpi.so.12 ABI.

Optional build flags:

  • Trims. Novomodelo uses no Fortran, C++, or MPI-IO bindings, so --disable-fortran --disable-cxx --disable-romio --disable-doc shorten the build without affecting Novomodelo.
  • Faster runtime. --enable-fast=O3,ndebug builds an optimised MPICH without internal assertions.

On a SLURM cluster, srun is the launcher: it places the ranks, binds them to cores, and (via PMIx) initialises MPI. Because --comm-backend defaults to auto, Novomodelo detects the srun launch and selects the MPI backend automatically.

The commands below use three placeholders. <partition> is the SLURM partition, <shared-path> is a directory every node can read that holds novomodelo-mpi and the batch script, and <case-dir> is the case directory on a shared filesystem, at the same path on every node because every rank reads it.

Novomodelo uses MPI for inter-node communication and rayon threads for intra-node LP solves. The launch below runs one rank per node, with the thread pool sized as set out in Sizing and scaling:

Terminal window
srun --mpi=pmix --nodes=4 --ntasks-per-node=1 --cpus-per-task=96 --exclusive \
./novomodelo-mpi run <case-dir> --threads 96

In this command --threads 96 matches --cpus-per-task=96, so each rank’s rayon pool has one thread per allocated core. --exclusive gives each rank a whole node with no co-tenants competing for cores or memory bandwidth. The 4 nodes and the 96 cores are illustrative: set them to the allocation your job needs.

A production job wraps that launch in an sbatch script. Replace the placeholders and set the node count, cores per task and time limit for your workload:

#!/bin/bash
#SBATCH --job-name=novomodelo-mpi
#SBATCH --nodes=4
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=96
#SBATCH --exclusive
#SBATCH --output=novomodelo_%j.out
#SBATCH --error=novomodelo_%j.err
#SBATCH --partition=<partition>
#SBATCH --time=0-08:00:00
MPICH_PREFIX=/opt/mpich
NOVOMODELO_MPI="<shared-path>/novomodelo-mpi"
export PATH="$MPICH_PREFIX/bin:$PATH"
export LD_LIBRARY_PATH="$MPICH_PREFIX/lib:${LD_LIBRARY_PATH:-}"
srun --mpi=pmix "$NOVOMODELO_MPI" run "$1" --threads "$SLURM_CPUS_PER_TASK"

Refer to the binary by an absolute path. sbatch runs the script in the directory you submitted from, so ./novomodelo-mpi works only if you happen to run sbatch from the directory that holds the binary.

Save the script as run-novomodelo.sh in <shared-path>, next to novomodelo-mpi. Submit it with the case directory as its argument:

Terminal window
sbatch <shared-path>/run-novomodelo.sh <case-dir>
FlagPurpose
--mpi=pmixPMIx process startup (recommended; see the tip below)
--mpi=pmi2PMI-2 startup. It needs an MPICH built with PMI-2 support; the --with-pmi=pmix build above has none
--nodes=NNumber of nodes
--ntasks-per-node=NMPI ranks per node (1 for the one-rank-per-node mapping)
--cpus-per-task=TCores per rank — pass the same value to --threads
--exclusiveGive each rank a whole node
--cpu-bind=coresPin each rank’s threads to specific cores
--mem-bind=localAllocate memory from the NUMA node closest to the bound cores

No measured scaling data is published, so these rules relate ranks, threads, forward passes and memory to one another instead of giving node or core counts; Performance Accelerators — Parallel Execution Model holds the facts behind them.

  1. Run one rank per node, with --threads equal to the cores allocated to the rank (--cpus-per-task). Each LP solve runs on one thread, so a rank solves at most --threads LPs at a time.
  2. Run every rank with the same --threads. Under a sampled forward pass, the run stops with an error when ranks differ in --threads.
  3. Under the sampled method, set training.selection.forward_passes and simulation.selection.num_scenarios to at least the rank count times --threads, and to a multiple of it for equal shares. On more than one rank with --threads above 1, a forward_passes that is not a multiple of the rank count stops one rank with an index out of bounds error while the other ranks wait, so the job runs until its wall time. Trajectories and scenarios are split over the ranks and then over each rank’s workers, and a worker left without work idles.
  4. Size each node’s memory for one rank’s full copy of the case, the opening tree, the cut pool and the stage templates. Each rank holds its own copy, so each extra rank on a node adds another, and the cut pool scales with the iteration budget and the forward passes.
  5. Before adding ranks or threads, read the printed Time split and the knob table in Performance Accelerators — Diagnosing Performance. The lower-bound evaluation runs serially on rank 0, and under a sampled forward pass every rank waits on the per-stage cut exchange, so neither shrinks with more ranks or threads.
  6. Place the case directory as the <case-dir> definition under Running under SLURM requires.

A job writes its checkpoint when training ends and, with periodic checkpoints on, on each scheduled iteration. A job killed at its --time loses the iterations after its last checkpoint, and the next job resumes from that checkpoint. An interrupted write that replaces a checkpoint leaves a complete one, old or new; an interrupted first write can leave none (see Resume).

To fit a long training into wall-time-limited jobs, split it into slices, one batch job each, and submit the next slice after the previous one ends. Every slice after the first sets policy.mode to "resume" and raises iteration_limit to the new total, which counts the iterations of all earlier slices (Train in Slices gives the rule sets and how to size a slice).

The job’s --time covers the case load, the load of the previous checkpoint, the slice’s iterations and the checkpoint write, plus the simulation and its result export in a slice that enables simulation (simulation.enabled set to true; the default is off). The checkpoint is written before simulation starts, so a job killed during simulation keeps that slice’s checkpoint.

After editing <case-dir>/config.json as above, submit each slice on the same <case-dir>. The batch script passes no --output, so every slice uses <case-dir>/output on the shared filesystem, and every rank loads the previous checkpoint from it:

Terminal window
sbatch <shared-path>/run-novomodelo.sh <case-dir>

AWS ParallelCluster reference architecture

Section titled “AWS ParallelCluster reference architecture”

Novomodelo’s production deployment runs on AWS ParallelCluster: a lightweight, always-on head node for submission plus one or more compute queues whose nodes are provisioned on demand and torn down when idle. Each queue is an independent SLURM partition that scales on its own.

Head node · light, always-onsubmit + SLURM controllerCompute queue (SLURM partition)N × compute nodes · scaled on demand1 rank × T threads per nodeFSx · shared filesystemcases · checkpoints · results sbatch → srun (PMIx)stage casecheckpoint + export

The moving parts:

  • Custom machine image (shared by head and compute nodes). MPICH is built from source into the image at /opt/mpich, using the ch4:ofi + EFA recipe above. Baking it into the image means every dynamically-launched compute node already has the runtime — no per-job install step.
  • The binary. The pre-built novomodelo-mpi archive is downloaded from the release page to <shared-path>. It is the same binary produced by the release CI, and the custom image’s system libraries are kept compatible with that build.
  • Interconnect (EFA). Because MPICH is built --with-libfabric=/opt/amazon/efa and --with-device=ch4:ofi, the OFI netmod selects the EFA provider on EFA-enabled instances automatically. For fabric-level diagnostics or provider overrides, consult the AWS EFA and libfabric documentation.
  • Instances. The compute queue in the example uses c7a.48xlarge nodes with the one-rank-per-node × --threads=$SLURM_CPUS_PER_TASK mapping; smaller compute-optimised instances work the same way.
  • Storage (FSx). The case directory, checkpoints, and exported results live on the shared FSx filesystem. Novomodelo’s per-solve hot path holds its working set in memory (its allocation discipline is documented in the novomodelo-sddp README); it touches the shared filesystem heavily only at checkpointing and result export, so FSx throughput is sized for those phases rather than the solve loop.

Submission is the same sbatch <shared-path>/run-novomodelo.sh <case-dir> shown above: the head node queues the job, ParallelCluster powers up the compute nodes in the target partition, srun --mpi=pmix launches one rank per node, and the nodes scale back down when the queue drains.


error while loading shared libraries: libmpi.so.12 — the MPI runtime is not on the loader path. Export LD_LIBRARY_PATH=/opt/mpich/lib:$LD_LIBRARY_PATH (or module load mpich) before launching, as the batch script does.

comm: local when you expected comm: mpi — either the binary was built without the mpi feature (check novomodelo version), or the process was not launched under a recognised launcher. Force the backend with --comm-backend mpi; it fails with a clear message on a non-MPI binary rather than silently running single-process.

Every rank reports Layout: 1 rank, and the run repeats once per node. Symptoms: the Execution banner appears once per rank, not once in total, and each copy reads Layout: 1 rank on <host>. The simulation reports ... across 1 ranks, and the separate copies collide on the shared output directory with write errors. MPI_Init fell back to a singleton: every process launched by srun became its own one-rank MPI_COMM_WORLD.

novomodelo-mpi version still reports comm: mpi and the banner still shows Backend: MPI, because the binary is fine. The fault is in the MPICH runtime: it was built without --with-pmix, or against a different PMIx than SLURM’s plugin, so the PMIx handshake failed silently.

To fix it:

  1. Confirm the cause with ldd <mpich-prefix>/lib/libmpi.so.12 | grep libpmix. Seeing no line, or a path outside the cluster’s PMIx, confirms it.
  2. Rebuild MPICH with --with-pmix pointing at the PMIx that SLURM uses (see Building MPICH from source).
  3. Point the batch script’s exports at the new build. novomodelo-mpi itself needs no rebuild.

Switching to --mpi=pmi2 does not fix this.

srun --mpi=pmi2 fails with pmijobid missing in fullinit command. An MPICH built --with-pmi=pmix (the recommended build) contains no PMI-2 client, so --mpi=pmi2 is not an alternative launch path for it. Use --mpi=pmix. The same error is also caused by an incompatibility between SLURM 24.05’s PMI2 plugin and MPICH’s PMI2 wire protocol. If only PMI2 is available and you built Hydra (i.e. not --with-pm=no), mpiexec bypasses PMI2 entirely.

Multi-node job hangs at MPI_Init on MPICH 4.3.x — the 4.3.0 PMI2 client regression described above. Downgrade to MPICH 4.2.x.


For one novomodelo-mpi binary, with every node running the same system image on one processor model (the same libm, libstdc++ and MPI library), the numerical content of the output files is bit-identical at any rank count and at any thread count (the same --threads on every rank); timestamps, wall-clock timing columns and the recorded execution topology differ. The guarantee holds for a stopping-rule set without time_limit: a wall-clock or signal stop depends on timing (time_limit). Validate on one rank with novomodelo-mpi itself, because the single-process novomodelo is a different build whose results are not promised equal. The full statement is in Determinism & Provenance.

Check your setup in three steps. Each step catches a different failure.

  1. The binary is MPI-enabled. This only checks the build, not the runtime wiring:

    Terminal window
    <shared-path>/novomodelo-mpi version # → comm: mpi
  2. MPICH links the cluster’s PMIx. Run the ldd … | grep libpmix check from Building MPICH from source.

  3. The ranks form one job. Run a two-node smoke test, which exercises both PMIx wire-up and the inter-node fabric:

    Terminal window
    srun --mpi=pmix --nodes=2 --ntasks-per-node=1 \
    <shared-path>/novomodelo-mpi run <case-dir> --threads 2

    Only rank 0 prints the Execution banner, so it must appear once. Its Backend and Layout lines on a multi-node job have this format, followed by one host line per node ({rank_word} is rank or ranks; {range} lists the host’s ranks, such as 0–1):

    Backend: MPI ({library_version}, {standard_version})
    Layout: {world_size} {rank_word} across {num_hosts} nodes
    {hostname}: ranks {range} ({count} {rank_count_word})

    If the banner appears twice, each copy reading Layout: 1 rank on <host>, the ranks fell back to singletons. See Troubleshooting.