Policy Management
Novomodelo stores the trained future-cost function (cuts) and LP basis in a policy
directory, together with the visited states when
exports.states is true. The policy
section of config.json controls where that directory lives and whether
training starts from scratch or from a prior checkpoint.
Policy Modes
Section titled “Policy Modes”The policy.mode field selects one of three initialization strategies. The
default is "fresh".
policy.path (default ./policy) resolves against the run’s output directory,
<case>/output unless novomodelo run --output names another (an absolute path is
used as is), so the default checkpoint is <case>/output/policy. Training writes
it there; warm-start, resume and simulation-only read it there.
policy.boundary.path differs: it resolves against the case
directory.
Fresh (Default)
Section titled “Fresh (Default)”Training starts from an empty future-cost function. All prior cuts in
policy.path are ignored (or the directory does not yet exist).
{ "policy": { "mode": "fresh" } }Use "fresh" for new studies or when you want a clean training run with no
influence from earlier iterations.
Warm Start
Section titled “Warm Start”Novomodelo loads the cuts from an existing policy checkpoint before training begins. Training then continues, adding new cuts on top of the loaded ones. The loaded cuts count as the initial future-cost approximation.
{ "policy": { "mode": "warm_start", "path": "./policy" } }Use "warm_start" to continue training from a previous run’s cuts, for example
with more iterations or other stopping rules. The loop runs up to the largest
iteration_limit in training.stopping_rules, and that limit counts only the new
iterations: a warm start with limit 20 runs up to 20 iterations on top of the
loaded cuts, whereas under Resume the limit also counts the completed
iterations. A warm-start load reads the whole saved policy, so it must pass every
check of the Policy Load Contract.
Resume
Section titled “Resume”Resume continues training from the checkpoint at policy.path (the one written
when training ended, or the latest periodic checkpoint; see
Checkpointing Configuration),
whichever stopping rule ended the earlier run. Novomodelo reads the checkpoint’s
completed_iterations count and the loop continues from the next iteration. The
load is a full-FCF load, so the checkpoint must pass every check of the
Policy Load Contract.
{ "policy": { "mode": "resume", "path": "./policy" } }Training writes a checkpoint when it ends, so resume needs no checkpointing
key; with periodic checkpoints on, it also writes one on each scheduled
iteration, and each replaces the last. A job killed before the end-of-training
write, for example by a scheduler’s wall-time limit, keeps the latest checkpoint
written before: the latest periodic one, or the previous run’s.
An interrupted write that replaces a checkpoint leaves a complete one, old or new;
an interrupted first write can leave none (see
Checkpoint Directory Contents).
Train in Slices fits a long training into wall-time-limited
jobs.
The loop runs up to the largest iteration_limit in training.stopping_rules.
That limit counts all iterations, the completed ones included; when the
checkpoint has already reached it, Novomodelo warns and trains no further.
The seed comes from config.json, not from the checkpoint, and each iteration’s
forward-pass noise derives from the seed and the iteration number. With the
training run’s seed, a resumed iteration therefore samples the same forward-pass
noise as that iteration in a run without a resume.
Resume restores the cuts, the stored LP bases that fit, the iteration count, and the lower bound
of every completed iteration.
Everything else comes from the current run: the seed, the stopping rules, and
the wall-clock timer that time_limit reads, which starts when the resumed
training starts.
Simulation-Only Mode
Section titled “Simulation-Only Mode”With training.enabled set to false and simulation enabled, novomodelo run skips
training, loads the checkpoint at policy.path, and simulates it.
policy.mode is ignored in this mode.
The load is a full-FCF load, so the checkpoint must pass every check of the Policy Load Contract: it comes from the running novomodelo version.
Start from the config.json that trained the policy and change only
training.enabled and the simulation settings. training.selection and
training.stopping_rules stay required even though training does not run, and
every key that shapes the stage LP must equal the training run’s (a stored basis
whose stage LP changed is left out, see Stored-basis gate). For
example, a policy trained with modeling.inflow_non_negativity set to
{ "method": "none" } loads without its stored bases if the simulation-only config omits the key:
it then defaults to "penalty", which adds slack columns (one per hydro) to the
stage LP.
The keys to change, merged into that config.json (all other keys stay as in
training):
{ "training": { "enabled": false, "selection": { "method": "sampled", "forward_passes": 1 }, "stopping_rules": [{ "type": "iteration_limit", "limit": 128 }] }, "simulation": { "enabled": true, "selection": { "method": "sampled", "num_scenarios": 100 } }, "policy": { "path": "./policy" }}With these keys in config.json, novomodelo validate accepts the case.
Validate reads the checkpoint at policy.path and applies the same load, so
this confirms the config and the policy load:
Valid case: 1 buses, 1 hydros, 2 thermals, 0 linesUse this mode to run additional simulation scenarios on a policy that has already converged, or to compare multiple saved policies on the same scenarios.
Train in Slices
Section titled “Train in Slices”Split a training into slices when a scheduler’s wall-time limit is shorter than
the training needs. Each slice is a separate job that resumes from the
previous slice’s checkpoint and trains up to a larger iteration total. A slice
must finish, and write its checkpoint, before its wall time ends: a job killed
first keeps only the iterations up to its latest checkpoint: the latest periodic
one with policy.checkpointing.enabled, or the previous slice’s. Run every slice on the
same case with the same output directory: each slice reads the previous
checkpoint from policy.path and writes its own checkpoint back to it. Under
MPI every rank loads that checkpoint, so the output directory is visible at the
same path on every node.
Slicing by iteration_limit works on any rank count. Size each slice from
the time_total_ms trend
in the previous slice’s training/convergence.parquet, which holds only the
iterations of the slice that wrote it. Per-iteration time grows as cuts
accumulate, so leave a margin.
Slice 1 trains to the first iteration total:
{ "training": { "stopping_rules": [{ "type": "iteration_limit", "limit": 50 }] }}Every later slice sets policy.mode to "resume" and raises limit to the new
total, which counts the iterations of all earlier slices:
{ "training": { "stopping_rules": [{ "type": "iteration_limit", "limit": 100 }] }, "policy": { "mode": "resume" }}A slice may end on a time_limit rule, measured
from the start of that slice’s training. Keep an iteration_limit holding the
total in the set (every rule set needs one), and leave
stopping_mode at its default, "any",
so that time_limit alone ends a slice. Novomodelo checks time_limit after each
iteration, so a slice runs up to one iteration past seconds. Set seconds
below the job’s wall time by that iteration, the case load, and the checkpoint
write. The same set serves every slice, with policy.mode set to
"resume" from the second slice on. This set ends each slice after 50 minutes of
a one-hour job and stops the training at 200 iterations:
{ "training": { "stopping_rules": [ { "type": "time_limit", "seconds": 3000 }, { "type": "iteration_limit", "limit": 200 } ] }}Checkpointing Configuration
Section titled “Checkpointing Configuration”With policy.checkpointing.enabled set to true, training writes a checkpoint
to policy.path after each scheduled iteration that does not end training:
initial_iteration (default interval_iterations), then every
interval_iterations, in absolute iteration numbers. The iteration that ends
training writes only the end-of-training checkpoint. Each checkpoint replaces
the previous one and holds the same files as the end-of-training one, with
completed_iterations set to its iteration. Rank 0 writes it. A failed write
ends training on every rank with checkpoint write at iteration {n} failed: {e}.
store_basis and compress have no effect; the keys are listed under
Configuration.
Checkpoint Directory Contents
Section titled “Checkpoint Directory Contents”A written checkpoint has the following layout under policy.path (resolved
against the output directory, as described in Policy Modes);
policy/ below stands for that directory:
policy/ manifest.bin -- study-global manifest: study graph, stage count, provenance + a format_version marker (FlatBuffers, written last as the commit signal) cuts/ 000.bin -- cut coefficients and intercepts for pool 0 (each self-describing its own cost-scale and graph identity) 001.bin -- cut coefficients and intercepts for pool 1 ... basis/ 000.bin -- LP basis for stage 0 (bases are always written with the terminal checkpoint) 001.bin ... states/ -- written only when exports.states is true 000.bin -- visited states, stage 0 001.bin ...File names are zero-padded to three digits (NNN.bin). Under cuts/, the id
is the pool id; under basis/, it is the policy-graph node’s canonical
0-based position (not its declared node id); under states/, it is the 0-based
stage position. On a stage chain the node position and the stage position
coincide.
The file name itself is not read for identity — each buffer carries its own
id internally, and the reader derives pool/stage identity from the payload,
never from the name.
manifest.bin is written last. Every write is staged in
<policy.path>.staging; the commit renames an existing <policy.path> to
<policy.path>.previous, renames the staged copy into place, then removes
.previous. Every reader (novomodelo run, novomodelo validate, the boundary load and
Python load_policy) uses the first of <policy.path>, .staging and
.previous that holds a manifest.bin.
An interrupted write that replaces a checkpoint leaves a complete one, old or new;
an interrupted first write can leave none. A load refuses a directory with no
manifest.bin, in itself or in either sibling. Before staging, a write refuses a
policy.path, or a .staging or .previous sibling, that is not a directory or
that holds an entry a checkpoint write does not leave there, and changes nothing
on disk; a checkpoint holds only manifest.bin, manifest.bin.tmp and
metadata.json at its top level, and only *.bin and *.bin.tmp files in its
cuts/, basis/ and states/ directories. The message is
refusing to write a checkpoint: {entry}, found in {dir}, is not part of a checkpoint (manifest.bin, metadata.json, cuts/, basis/, states/); move it elsewhere or write the checkpoint to another directory.
A symbolic link at policy.path is followed once: the commit runs at its target,
the siblings are the target’s, and the link is kept.
manifest.bin is read first: its format_version marker is checked before
any cut, basis, or state payload is parsed (step 1 of the
check order).
manifest.bin records the study graph, the number of stages, the producing
novomodelo version, and the producer provenance (completed iterations, lower- and
upper-bound values, forward passes per iteration, and the RNG seed) — see
Policy Checkpoint for the full
field-by-field table. Whether a checkpoint loads into the current study is
decided by the Policy Load Contract.
Policy Load Contract
Section titled “Policy Load Contract”Warm-start, resume, and simulation-only runs are full-FCF loads: they read
every cut pool of the future-cost function (FCF), the study graph, and the
stored LP bases from the checkpoint at policy.path. The checks below run in
order; a refusal at steps 1–4 stops the run before training or simulation
starts, and step 5 leaves out the stored bases that do not fit.
A boundary load runs its own checks, listed under
Compatibility requirements.
Check order
Section titled “Check order”- Read.
manifest.binis present and its wireformat_versionis3(theformat_versiongate). Amanifest.binstamped with another value is refused with:unsupported checkpoint manifest format_version {format_version}; expected 3; re-run the program that produced it with novomodelo {running}; for a converted boundary policy, convert it again - Cost scale. The terminal pool records its
cost_scale_factor(see Cost-Scale Canonicalization). - Software and version. The recorded
softwareandsoftware_versionequal the running software’s name and version exactly (see Version gate). - State and graph. The state dimension, stage count, pool count, per-slot entity identity, and study graph (nodes and edges) equal the current study’s.
- Stored bases. A stored LP basis that does not fit its stage LP is left out with one warning, and the load proceeds (see Stored-basis gate).
A refusal at step 4 means the study differs from the one that trained the policy. The message names the mismatch; restore the matching inputs, or train a new policy for the changed study.
Version gate
Section titled “Version gate”A policy loads only in the software and version that wrote it. Every load path —
warm-start, resume, simulation-only, and boundary, from the CLI and from Python —
requires the software and software_version recorded in manifest.bin to
equal the running software’s name and version exactly. An older or newer version is refused, and so is a different
patch or pre-release build (for example 0.17.0 against 0.17.1).
The refusal reads:
policy was written by {writer}, but this is novomodelo {running}; a policy loads only in the software and version that wrote it: re-run the program that produced it with novomodelo {running}; for a converted boundary policy, convert it again{writer} names the recorded software and version: <software> <version>,
<software>, which recorded no version, software that recorded no name, version <version>,
or software that recorded no name or version; {running} is the version of
the novomodelo that is loading it. novomodelo run reports
the refusal as a validation error before training or simulation starts;
Error Codes lists the exit code
and the Python exception.
To load the policy, retrain it with the running version. Alternatively,
re-export it by re-running the pipeline that wrote the checkpoint under the
running version: the script that calls
novomodelo.write_policy_checkpoint,
or the boundary import of Case Conversion (novomodelo-bridge).
A manifest that parses (step 1) is not thereby accepted: this gate decides whether the running version may load it. FlatBuffers Policy Schema — Versioning policy covers the wire level.
Stored-basis gate
Section titled “Stored-basis gate”A full-FCF load uses a stored LP basis only when it fits its stage LP, leaves
the others out with one warning, and proceeds. For example, novomodelo run prints:
warning: stored bases not used: 4 of 4 do not fit the current LP (first: node 0, 28 columns, the LP has 29); a stored basis is used only when its column count equals the LP's, its row count equals the LP's template rows plus its recorded cut rows, and its basic count equals its row count; the policy was trained on a different LPThe node and the counts come from an example case and differ per study.
A stored LP basis fits when its column count equals its stage LP’s, its row count equals that LP’s template rows (the rows before any cut) plus the cut rows the basis itself records, and its count of basic entries equals its row count.
Warm-start and resume check the bases a checkpoint carries, simulation-only always checks them, and a boundary load checks none.
Every training run writes its stage LP bases with the end-of-training
checkpoint, and policy.checkpointing.store_basis has no effect on that. Any
change that alters the column count of a stage LP therefore leaves out the
stored bases of the stages it changes: in training those nodes start their first
solve cold and then reuse the bases they capture, and in a simulation-only run
they solve without a stored basis.
Examples are a different block count, a different entity set, and a different
block mode on a stage with hydro plants and more than one block.
novomodelo validate runs the same load
and reports the same warning.
Cost-Scale Canonicalization
Section titled “Cost-Scale Canonicalization”Novomodelo’s objective cost-scale factor
(modeling.cost_scale_factor)
is configurable per study. Cut coefficients and intercepts are computed in the
solving study’s internally scaled cost space, so policy export and load
convert between that scaled space and a canonical, scale-independent
representation at rest:
- Export multiplies every cut coefficient and intercept by the
writing study’s
cost_scale_factor, so a persisted policy holds canonical currency units — not the writer’s internal scaled cost space. - Every load path — warm-start, resume, simulation-only, and boundary-cut
injection — divides by the loading study’s own
cost_scale_factor, even when it equals the writer’s, converting the canonical values back into that study’s internal scaled space.
Each cuts/<pool>.bin (see the field table in
Policy Checkpoint) carries its own
cost_scale_factor provenance field recording the writing study’s factor — the
checkpoint is self-describing per pool. Every load path — warm-start, resume,
simulation-only, and boundary-cut injection — requires each resolved pool to
carry it: a pool whose cost_scale_factor reads absent is rejected. A full-FCF
load refuses with
policy checkpoint predates self-describing cuts (its resolved cuts/<pool>.bin carries no cost_scale_factor); re-run the program that produced it with novomodelo {running}; for a converted boundary policy, convert it againA boundary load refuses with its own message, quoted in item 3 of the boundary load order.
Loading applies one floating-point division to every cut coefficient and intercept, moving each value by up to a few units in the last place (ULP). This is below solver tolerance, but a bit-exact hash over a loaded policy can differ from one over the in-memory policy that wrote it.
Boundary Cuts
Section titled “Boundary Cuts”Boundary cuts allow a Novomodelo study to load terminal-stage future cost function (FCF) approximations from a different Novomodelo policy checkpoint. This is the mechanism for chained studies: a downstream study imports one cut pool of an upstream study’s policy as its terminal boundary condition, so its end-of-horizon decisions account for the upstream study’s future cost of water. The downstream study may have a shorter horizon or finer stages than the upstream one. For the method, see Post-Study Boundary & Chained Studies §5.
How it works
Section titled “How it works”- Run the upstream study to produce a policy checkpoint, or author the checkpoint in Python (Read and Write Checkpoints from Python).
- Run the downstream study with
policy.boundarypointing to the upstream checkpoint. Novomodelo selects the source pool priced at the downstream study’s terminal boundary date — the last non-negative stage’send_date— and injects its cuts into the terminal stage’s row pool as fixed boundary conditions. There is no stage-index knob: the pool is chosen by matching that date against each source pool’s own priced-state date.
The imported boundary cuts are not updated by the SDDP training algorithm.
Their coefficients remain fixed throughout training and simulation, and the
active ones provide a floor on the terminal-stage future cost.
The training.cut_selection.max_active_per_stage setting
(cut_selection) caps the active cuts
of every pool, including the terminal pool that holds the imported boundary
cuts, and only cuts generated in the current iteration are exempt, so a cap
below the number of active imported boundary cuts evicts (deactivates) some of
them.
Configuration
Section titled “Configuration”Add a boundary object to the policy section of config.json:
{ "policy": { "mode": "fresh", "boundary": { "path": "../upstream_study/output/policy", "strict": false } }}| Field | Type | Default | Description |
|---|---|---|---|
path | string | — | Path to the source checkpoint directory; see policy.boundary for how it resolves. |
strict | boolean | false | Governs a superset source (one pricing an entity or commitment this study does not model). false drops the surplus, records it in the reconciliation report, and loads; true rejects the load. |
When boundary is absent or null, no boundary cuts are loaded (the default).
Compatibility requirements
Section titled “Compatibility requirements”Unlike a full-FCF load — which requires the saved
policy’s state dimension and per-slot entity layout to match the current study
exactly (step 4 of the check order, a check that cannot be
disabled) — boundary injection tolerates a source of a different state
shape. The source’s terminal state is reconciled onto the current study’s own
state per slot by entity identity and its per-family slot dates (the
inflow-lag reference_date, the forward-family interval_start/interval_end),
not by position: a source trained with a different
set of state coordinates — no in-transit buckets, or monthly anticipated slots
feeding a differently-shaped study — still injects, each source coefficient
binding to the current study’s slot for the same entity. A source coordinate
that couples a state slot the current study does not model cannot be carried; it
is reported in a per-family reconciliation summary and dropped, and the load
still succeeds by default. See
Post-Study Boundary & Chained Studies §4 for the
reconciliation mechanism.
That leniency is the default reconcile path, but the load still enforces hard requirements. A boundary load runs its checks in this order, and the first failure stops the run:
- Read. The checkpoint at
policy.boundary.pathcan be read; a read failure is refused withfailed to read boundary policy checkpoint at {path}: {e}, where{e}is the underlying error. - Pool selection. Exactly one pool of the source is priced at the study’s terminal boundary date.
- Cost scale. The resolved pool records its
cost_scale_factor(see Cost-Scale Canonicalization); a pool without it is refused withboundary policy checkpoint at {path} predates self-describing cuts (its resolved cuts/<pool>.bin carries no cost_scale_factor); re-run the program that produced it with novomodelo {running}; for a converted boundary policy, convert it again. - Single-node source. The resolved pool is a single-node terminal pool, not
one shared across several study nodes. A shared pool is refused with
boundary policy at {path}: a boundary source must be a single-node terminal pool; the resolved pool is shared by multiple nodes. - Inflow-lag depth. Before building the study, Novomodelo reads the deepest
inflow-lag slot that any cut pool of the boundary checkpoint references and
extends the study’s inflow-lag state to at least that depth
(Post-Study Boundary & Chained Studies §5.2);
when the depth is above zero,
novomodelo runwithout--quietprintsBoundary policy: inflow-lag depth {d}while loading. This check confirms that the resolved pool references no lag deeper than the extended state, so a user case does not trigger it. A message startinginternal:is a Novomodelo defect; report it. - Season and PAR order. When the study declares a season cycle, the
source’s
season_manifestis present and agrees with the study’s season cycle, season count, and per-hydro PAR orders. - Topology subset. When both checkpoints carry an entity manifest, the source prices every hydro storage slot and every inflow-lag slot of the current study.
- Software and version. The software and version recorded in the source checkpoint equal the running software’s name and version exactly (see Version gate); the checks above run first, so a source written by another software or version can fail one of them before this one.
- State dimension. If either checkpoint has no entity manifest, slots cannot
be matched by identity, so the topology-subset and slot-date checks are
skipped and the source pool’s state dimension must equal the study’s. The
refusal starts
boundary policy state_dimension mismatch: policy has {n}, current system has {n}. - Slot dates. Every dated in-transit or anticipated slot of the source has
an
interval_startandinterval_endthat decode to a date range of positive length. A slot that fails is refused with one of:boundary policy source forward-family slot (entity_id={entity_id}, subindex={subindex}) has an undecodable interval_start {interval_start}boundary policy source forward-family slot (entity_id={entity_id}, subindex={subindex}) has an undecodable interval_end {interval_end}boundary policy source forward-family slot (entity_id={entity_id}, subindex={subindex}) has a non-positive delivery interval: interval_start={interval_start}, interval_end={interval_end}
- Superset. Applies only with
strict: true: a source that prices more than the study models is refused (withstrict: falsethe surplus is dropped and reported).
Every check except Superset is independent of strict.
Date-selection rejects. Before any coefficient is reconciled, Novomodelo resolves the source to exactly one pool priced at the study’s terminal boundary date (list item 2). Three conditions abort the load with a named error:
- Undated source — no pool in the source checkpoint records a priced-state
date:
boundary policy checkpoint at {path} carries no priced_state_date on any pool (a pool written before priced dates were recorded); re-run the program that produced it with novomodelo {running}; for a converted boundary policy, convert it again. - No pool at the date — no source pool is priced at the boundary date:
boundary policy at {path}: no pool is priced at the study's boundary date {date} (available: {dates}). - More than one pool at the date — the source is ambiguous:
boundary policy at {path}: more than one pool is priced at the study's boundary date {date} (pools {pools}); boundary injection requires a unique priced source.
Season / PAR-order rejects. When the loading study models a seasonal or
per-hydro autoregressive (PAR) inflow process, the boundary source’s season
descriptor (season_manifest) must agree with the study’s. This check is
always-on — it is not gated by strict, and fires whenever the loading
study declares a season cycle (a study modelling no seasonal/PAR inflow skips it
entirely). It aborts the load with a validation error on the boundary-load
path — distinct from the version gate refusal. A boundary
source is rejected when its season_manifest:
- is absent — the source predates the season descriptor:
boundary policy checkpoint at {path} predates the season descriptor (its manifest carries no season cycle or PAR orders); re-run the program that produced it with novomodelo {running}; for a converted boundary policy, convert it again. - disagrees on the season cycle —
boundary policy at {path}: season cycle mismatch (study is {monthly|weekly|custom}, source is {…}). - disagrees on the season count —
boundary policy at {path}: season count mismatch (study has {n} seasons, source has {n}). - omits a hydro the study models —
boundary policy at {path}: hydro {id} has a modeled inflow season/PAR-order entry in the current study but none in the boundary source (the boundary was fitted on a different set of inflow processes). - disagrees on a hydro’s PAR-order count —
boundary policy at {path}: hydro {id} carries {n} PAR orders in the current study but {n} in the boundary source. - disagrees on a per-season PAR-order value (naming the hydro and the first
differing season) —
boundary policy at {path}: hydro {id} PAR order mismatch at season index {season} (the season's 0-based position in the cycle, not its id; study has order {n}, source has order {n}).
The season_manifest field shape is documented in
Policy Checkpoint.
Subset rejects. If the source prices less than this study needs, the load
is rejected regardless of strict:
- Missing storage — the source does not price a hydro the study models:
boundary policy does not price {names}; it was trained on a different set of plants. - Missing inflow-lag — the source has no inflow-lag coefficient at a depth
the study needs:
boundary policy has no inflow-lag coefficient for {names}: the boundary is lag-depth-incompatible with the current study.
Lag-date reject. A source inflow-lag coefficient and the current study’s
inflow-lag slot for the same hydro and depth may each carry a reference date.
When both are dated and the dates differ, the load is rejected regardless of
strict; this reject runs after list item 10 and before list item 11:
boundary policy's inflow-lag coefficient for hydro {id} at lag depth {depth} references a different past than the current study: boundary {date}, current {date}.
Lag coverage and dating. An interior source pool, any pool other than the
upstream study’s terminal pool, carries inflow-lag coefficients only when its
successor stage’s
stages[].state_variables
selection includes inflow_lags, which defaults to false. The terminal pool
of a trained upstream study is sized by that study’s full state, with every
inflow-lag slot of that state, whatever the selection; training never adds a
cut to it, so it holds only the cuts loaded into it, such as that study’s own
boundary cuts. The current (downstream) study’s terminal state has inflow-lag
slots when its own inflow model is autoregressive or when any pool of the
source references a lag (the extension of list item 5), and a source pool that
lacks a coefficient for any of those slots is rejected with the
Missing inflow-lag message. An autoregressive upstream study always
triggers that extension, because its terminal pool references every lag of its
state even when it holds no cuts. To chain from an interior pool, set
inflow_lags: true on the upstream stage that follows the source pool’s stage.
Each study dates its own lags: lag 1 by the start date of the pool’s own stage,
and each deeper lag by the start date of one stage further back in that
study’s full stage list, pre-study stages included. The source pool’s stage and
the current study’s last stage must therefore start on the same date, and so
must each pair of stages the same number of places before them, down to the
deepest lag both sides date, or the lag-date reject applies. In particular, it
applies wherever several stages of the current study together span one source
stage within that depth.
Superset behavior (strict). When the source prices more than the study
models, the surplus source slots are dropped during reconciliation either way.
With strict: false (the default) the drop is recorded in the reconciliation
report and the load proceeds; with strict: true the same condition rejects the
load: boundary policy at {path} is a superset: {summary}; see the reconciliation report or set policy.boundary.strict = false, where {summary} names every
dropping family and its count. strict governs only this superset case — it
does not relax the subset rejects or loosen the date-selection rejects.
Where the reconciliation report appears. When the load proceeds (a strict rejection shows only the superset message’s summary), the
per-family reconciliation is surfaced on three
paths:
novomodelo runprints a Boundary policy block to stderr while it loads the case, before training —Cuts loaded: N (priced at {date}),Source: {path}, and aReconciliation:lineN copied, N fanned out, N defaulted to 0.0, N dropped.novomodelo validate --jsonemits a boundary object withconfigured,boundary_date, andreportfields carrying the full per-family report (worked example under Configuration — boundary).novomodelo validate(human mode) printsboundary policy priced at {date}followed by the one-line summaryboundary reconciliation: N copied, N fanned out, N defaulted to 0.0, N source slots dropped.
Production coupling workflow
Section titled “Production coupling workflow”A chained-study pipeline uses boundary cuts as follows:
The upstream study writes a policy checkpoint whose cut pools are each priced at
a date. The downstream study names that checkpoint in policy.boundary.path,
and the pool priced at the downstream study’s terminal boundary date becomes its
terminal future cost function. The dependency runs one way: the upstream study
never reads the downstream one. The conditions under which the imported cuts
apply to the downstream study’s state are in
Post-Study Boundary & Chained Studies §5.2.
Interaction with warm-start
Section titled “Interaction with warm-start”Boundary cuts and warm-start are independent features. You can combine them:
{ "policy": { "mode": "warm_start", "path": "./policy", "boundary": { "path": "../upstream_study/output/policy" } }}This loads the downstream study’s own cuts from its previous run via warm-start AND loads the upstream policy’s boundary cuts at the terminal stage. Both sets of cuts contribute to the lower bound.
Read and Write Checkpoints from Python
Section titled “Read and Write Checkpoints from Python”novomodelo.results.load_policy reads a policy checkpoint into plain Python dicts,
and novomodelo.write_policy_checkpoint writes one from dicts of the same shape, so
a loaded checkpoint round-trips: load, edit, write. The supported use is
authoring an external boundary future-cost function, for example one derived
analytically or produced by another tool, as a checkpoint that Novomodelo loads
through policy.boundary. The on-disk AffinePiece layout is
documented in the FlatBuffers Policy Schema.
A study loads the checkpoint only in the novomodelo version that wrote it
(Version gate); novomodelo.results.load_policy does not check the version.
Load a policy
Section titled “Load a policy”novomodelo.results.load_policy(output_dir, policy_subdir="policy")output_dir is the study’s output directory; load_policy reads the
checkpoint from <output_dir>/<policy_subdir> (policy_subdir defaults to
"policy"). It returns a dict with three top-level keys:
| Key | Type | Description |
|---|---|---|
metadata | dict | format_version, software, software_version, created_at, num_stages, graph_manifest, season_manifest, producer |
stage_cuts | list[dict] | One entry per cut pool (one per stage on a stage chain) — see below |
stage_bases | list[dict] | One entry per saved LP basis: stage_id, iteration, column_status, row_status, num_cut_rows |
Each stage_cuts entry carries:
| Key | Type | Description |
|---|---|---|
stage_id | int | Pool id (0-based), the key that names cuts/<pool>.bin; the 0-based stage position on a stage chain |
state_dimension | int | Length every cut’s coefficients must have |
capacity | int | Cut-pool slot capacity |
warm_start_count | int | Cuts loaded from a previous artifact at run start |
populated_count | int | Number of populated cut slots |
cost_scale_factor | float | Cost-scale factor of the writing study, recorded for this pool (see Cost-Scale Canonicalization) |
node_id | int | Policy-graph id of the node that owns the pool; -1 for a pool with no single owning node |
graph_stage_id | int | Graph-stage id of the node that owns the pool; -1 when unresolved |
priced_state_date | int | The owning stage’s end_date encoded YYYYMMDD, the date the pool’s cuts price; the int32-minimum sentinel when not recorded. A boundary load selects its source pool by this date |
entity_manifest | list[dict] | Per-slot entity markers — see below |
cuts | list[dict] | Cut records — see below |
Each entry in cuts carries cut_id, slot_index, iteration,
forward_pass_index, intercept, coefficients, and is_active. The Python
dict key for a cut’s id is cut_id (the FlatBuffers wire field is
piece_id; cut_id is the record-level name used everywhere in this API).
entity_manifest is a list of per-slot markers,
{entity_type, entity_id, subindex, was_active, reference_date, interval_start, interval_end},
that let an externally authored boundary-cut checkpoint round-trip through
write_policy_checkpoint without losing slot identity. Which date marker a slot
carries follows from its family. A HydroInflowLag slot carries
reference_date, the start_date of its reference past stage, encoded
YYYYMMDD. The forward families (HydroTransitBucket,
AnticipatedThermalState) carry the half-open interval_start/interval_end
window, both encoded YYYYMMDD; an anticipated slot carries it only while it
holds a live commitment at the pool’s stage. A date field that does not apply
reads the int32-minimum sentinel (-2147483648): every other slot reads it in
all three fields, and so does an anticipated slot that holds no live commitment
at the pool’s stage.
load_policy applies no load gate: it reads any checkpoint this build can parse,
whichever novomodelo version wrote it, and returns the recorded software and version
as metadata["software"] and metadata["software_version"], without comparing
them with a study. A missing directory, or one without manifest.bin, raises
FileNotFoundError when neither <policy_subdir>.staging nor
<policy_subdir>.previous holds a manifest.bin, and an unreadable buffer
raises OutputError. Error Codes
lists the failure class of every load path.
Write a checkpoint
Section titled “Write a checkpoint”novomodelo.write_policy_checkpoint(path, stage_cuts, metadata, stage_bases=None, stage_states=None, inflow_lag_depth=None)stage_cuts and metadata mirror the load_policy shapes above, so a
checkpoint loaded from disk round-trips through Python without reshaping, with
one caveat: active_cut_indices is written by write_policy_checkpoint but not
returned by load_policy, so a load → write cycle resets it, and cut activity
round-trips through each cut’s is_active flag instead. stage_bases and
stage_states default to empty when omitted.
An authored metadata needs created_at, num_stages, and a producer with
completed_iterations, final_lower_bound, max_iterations, forward_passes,
warm_start_cuts, and rng_seed. format_version, graph_manifest,
season_manifest, and the other producer keys (best_upper_bound,
warm_start_counts, total_visited_states, training_block_mode,
training_block_mode_per_stage, cost_scale_factor, lower_bound_history (a
list of floats, empty when omitted)) are optional.
inflow_lag_depth, when set to N > 0, has Novomodelo reserve N canonical
HydroInflowLag state slots per storage hydro in every stage’s manifest and
place each cut’s inflow_lag_coefficients at their (hydro, depth) positions,
so a boundary policy authored for a case with no autoregressive inflow model can
still carry an inflow-lag-coupled terminal cost-to-go. Absent or 0, the
written checkpoint is byte-identical.
The function writes manifest.bin, cuts/, and basis/ directly under path,
and states/ only when stage_states is non-empty. That is the
Checkpoint Directory Contents layout, so
path should be a policy/-style directory. To use it as
policy.boundary.path, pass it relative to the consuming case directory, or as
an absolute path. load_policy joins policy_subdir onto the directory it is
given, so a checkpoint written to <dir>/policy is read back with load_policy
on its parent directory, <dir>.
write_policy_checkpoint stamps the running novomodelo version into the checkpoint
and ignores software, software_version or novomodelo_version keys in metadata, so the checkpoint loads only in
that version (see Version gate). A checkpoint written without
stage_bases is not subject to the stored-basis gate: it
carries no bases to check.
Round-trip example
Section titled “Round-trip example”The following example is copy-runnable. To edit a trained policy, load it with
novomodelo.results.load_policy("output/") and apply the edit step below to
loaded["stage_cuts"]; the example authors a checkpoint from scratch so it runs
without a study. It writes the checkpoint, loads it back, appends a third cut,
writes the edited policy back to the same path, and prints the reloaded cut
count and the novomodelo version recorded in the checkpoint.
import tempfile
import novomodeloimport novomodelo.results
parent = tempfile.mkdtemp()
stage_cuts = [ { "stage_id": 0, "state_dimension": 3, "capacity": 10, "cuts": [ { "cut_id": 1, "slot_index": 0, "iteration": 1, "forward_pass_index": 0, "intercept": 42.0, "coefficients": [1.0, 2.0, 3.0], "is_active": True, }, { "cut_id": 2, "slot_index": 1, "iteration": 1, "forward_pass_index": 1, "intercept": 10.5, "coefficients": [0.5, -1.5, 2.5], "is_active": True, }, ], }]
metadata = { "created_at": "2026-07-30T00:00:00Z", "num_stages": 1, "producer": { "completed_iterations": 5, "final_lower_bound": 123.45, "best_upper_bound": 130.0, "max_iterations": 10, "forward_passes": 4, "warm_start_cuts": 0, "warm_start_counts": [0], "rng_seed": 42, "total_visited_states": 0, "training_block_mode": "parallel", "training_block_mode_per_stage": [], "cost_scale_factor": 2_500_000.0, },}
# Write to `<parent>/policy` — `load_policy` joins policy_subdir="policy" onto# the directory it is given, so it is loaded from `parent` below.novomodelo.write_policy_checkpoint(f"{parent}/policy", stage_cuts, metadata)
loaded = novomodelo.results.load_policy(parent)assert loaded["metadata"]["format_version"] == 3assert loaded["stage_cuts"][0]["populated_count"] == 2
# Edit: append a third cut. Its coefficients must match state_dimension (3).# Keep cut_id and slot_index unique, and set populated_count to the new cut count.cuts = loaded["stage_cuts"][0]["cuts"]cuts.append( { "cut_id": 3, "slot_index": 2, "iteration": 2, "forward_pass_index": 0, "intercept": -5.0, "coefficients": [4.0, -2.0, 0.0], "is_active": True, })loaded["stage_cuts"][0]["populated_count"] = len(cuts)
# Write back: same round-trip shape, one more cut than before.novomodelo.write_policy_checkpoint( f"{parent}/policy", loaded["stage_cuts"], loaded["metadata"])
# The reloaded checkpoint holds three cuts and records the running novomodelo version.reloaded = novomodelo.results.load_policy(parent)print(len(reloaded["stage_cuts"][0]["cuts"]), reloaded["metadata"]["software_version"])The script prints the reloaded cut count and the version that
write_policy_checkpoint stamped:
3 0.18.0Input validation
Section titled “Input validation”write_policy_checkpoint raises ValueError when any of the following holds.
The message locates the problem: the stage and cut for the two per-cut checks,
the stage for a state-data mismatch, and the hydro for an unplaceable
inflow-lag coefficient.
- a cut’s
coefficientslength does not match its stage’sstate_dimension; - a stage’s state data length does not match
count * state_dimension; - a cut carries
inflow_lag_coefficientswithout a positiveinflow_lag_depth; - (under
inflow_lag_depth) a manifest lacks a leading storage block; - (under
inflow_lag_depth) an inflow-lag coefficient is unplaceable; or path(for a link, its target) or its.stagingor.previoussibling holds an entry no checkpoint writer leaves there (novomodelo.errors.ValidationError, aValueError, with therefusing to write a checkpoint: …text; nothing on disk changes).
Each snippet below uses metadata from the round-trip example. The second also
uses its stage_cuts; the third defines its own. Each snippet ends with the
exception text it raises.
A cut whose coefficients length disagrees with its stage’s
state_dimension:
bad_cuts = [ { "stage_id": 0, "state_dimension": 3, "capacity": 10, "cuts": [ { "cut_id": 1, "slot_index": 0, "iteration": 1, "forward_pass_index": 0, "intercept": 42.0, "coefficients": [1.0, 2.0], # length 2, but state_dimension=3 "is_active": True, }, ], }]
novomodelo.write_policy_checkpoint("output/policy", bad_cuts, metadata)# ValueError: stage 0 cut 1: coefficients has 2 entries, expected state_dimension=3A stage_states entry whose flat data length disagrees with
count * state_dimension:
stage_states = [ { "stage_id": 0, "state_dimension": 3, "count": 2, "data": [1.0, 2.0, 3.0], # length 3, but count * state_dimension = 6 }]
novomodelo.write_policy_checkpoint( "output/policy", stage_cuts, metadata, stage_states=stage_states)# ValueError: stage 0: states data has 3 entries, expected count*state_dimension=6A cut carrying inflow_lag_coefficients with no inflow_lag_depth passed to
reserve the lag slots (keyed by integer hydro id; list index 0 = lag
depth 1):
stage_cuts = [ { "stage_id": 0, "state_dimension": 3, "capacity": 10, "cuts": [ { "cut_id": 1, "slot_index": 0, "iteration": 1, "forward_pass_index": 0, "intercept": 42.0, "coefficients": [1.0, 2.0, 3.0], "is_active": True, "inflow_lag_coefficients": {0: [0.5]}, }, ], }]
novomodelo.write_policy_checkpoint("output/policy", stage_cuts, metadata)# ValueError: stage 0 cut 1: inflow_lag_coefficients supplied without inflow_lag_depth; pass inflow_lag_depth=N to reserve the lag slotsSee Also
Section titled “See Also”- Theory: Cut Management — the cut pool that a policy checkpoint persists across runs.
- Configuration — every
config.jsonfield documented - Running Studies — common workflows including training-only and simulation-only runs
- Output Format — detailed description of every output file