Stopping Rules
Purpose
Section titled “Purpose”This chapter defines the available stopping rules for the Novomodelo SDDP solver and how they combine. It covers iteration limits, time limits, bound stalling, and the recommended gap-based stopping criterion.
1 Available Stopping Rules
Section titled “1 Available Stopping Rules”SDDP can terminate based on multiple criteria. Each rule is evaluated independently, and the stopping_mode determines how they combine:
-
"any": Stop when any rule triggers (OR logic) -
"all": Stop when all rules other than the iteration limits trigger at the same iteration (AND logic); the largest iteration limit caps the run
2 Iteration Limit
Section titled “2 Iteration Limit”Evaluation:
where is the current iteration and is the limit.
Purpose: Safety bound to prevent infinite loops.
3 Time Limit
Section titled “3 Time Limit”Evaluation:
Wall-clock time is checked at the end of each iteration.
4 Bound Stalling
Section titled “4 Bound Stalling”Evaluation:
Track the lower bound over iterations. Once bounds have been recorded, compute the relative improvement over the window of the last recorded lower bounds, (the iterations parameter):
Stopping condition:
where is the bound-stalling tolerance on the relative improvement, set by the rule’s tolerance parameter.
Interpretation: The bound has plateaued — the relative improvement across the last recorded lower bounds is below the specified tolerance, indicating diminishing returns from further iterations.
5 Gap-Based Stopping (Recommended)
Section titled “5 Gap-Based Stopping (Recommended)”Illustrative bound evolution across iterations . In both panels the lower bound rises monotonically (the append-only cut pool). Left, a sampled forward pass: the upper bound is an estimate, and its confidence band reflects a sampling error that persists as the gap closes, so late in the run the band straddles the lower bound and certifies nothing. Right, an enumerated forward pass: the exact upper bound carries no band, stays above the lower bound and closes on it — the comparison the gap rule tests.
Optimality gap
Section titled “Optimality gap”At iteration , the optimality gap is the difference between the upper bound and the lower bound , whichever mechanism supplies the upper bound — the statistical estimate of a sampled forward pass or the exact bound of an enumerated one (Upper Bound Evaluation):
The percent gap normalizes it by the lower bound:
Both bounds are in unscaled currency units — the original cost units, with the cost scaling of the stage LP undone — so is in currency units and the percent gap is dimensionless. The denominator is the magnitude of the lower bound, never the upper bound, floored at one currency unit so that the ratio stays bounded when the lower bound is near zero. The per-iteration reported gap and the relative arm of the gap rule below share this denominator, so the two never disagree on what “relative gap” means.
The gap rule terminates training once the exact upper bound has closed on the lower bound.
Evaluation:
The rule evaluates the optimality gap with the exact upper bound, , and compares it clamped at zero, ; the clamp belongs to the rule’s comparison, not to the gap’s definition. In exact arithmetic the exact upper bound satisfies ; the clamp only absorbs floating-point noise once the gap has closed to (numerically) zero.
Stopping condition — two arms combined by disjunction (either is sufficient):
The absolute tolerance , in currency units, is set by the rule’s tolerance parameter. The relative tolerance , in percent, is set by the rule’s relative_tolerance parameter. Either tolerance may be configured alone; when both are configured, training stops as soon as either arm is satisfied.
Admissibility: the comparison above is only valid when the upper bound it uses is the exact bound, not a statistical estimate. Both of the following must hold:
- The forward pass enumerates the scenario tree exhaustively rather than sampling it, so the upper bound is exact — carrying no sampling error — rather than a sampled approximation.
- The risk measure is uniform across all stages: either expectation at every stage, or one CVaR measure (the same risk-aversion weight and tail fraction ) at every stage. Under expectation the exact bound is the probability-weighted cost over every scenario ; under a uniform CVaR it is the nested, time-consistent risk bound over the same enumerated tree (Upper Bound Evaluation §2). A stage-varying measure — the pair differing across stages, where every stage with is the expectation whatever its — is not admissible, because no single measure then aggregates the tree.
Under a sampled forward pass, or a stage-varying measure, the exact-bound comparison the rule depends on is unavailable, and the gap rule is rejected at setup with a named validation error rather than silently evaluating against an unsound comparison.
Why recommended: unlike bound stalling (a proxy for convergence) or the iteration and time limits (safety bounds unrelated to convergence), a closed gap is a direct optimality certificate under the exact-bound regime the rule requires — an enumerated forward pass with a risk measure held uniform across stages.
The admissibility decision takes the two conditions in order: first whether the forward pass is enumerated, then whether the measure is the same at every stage. Passing both admits the gap rule, which then compares the lower bound with the exact upper bound, so that a closed gap certifies optimality under the Tier 3 hypotheses; failing either rejects the rule at setup — stop such a run with bound stalling and an iteration limit instead, as in Sampled runs.
Sampled runs
Section titled “Sampled runs”Under a sampled forward pass the gap rule is not available. Stop a sampled run with bound stalling (§4) and an iteration limit (§2), combined so that the first to trigger ends the run (§7). The upper bound of a sampled run is a statistical estimate: its confidence interval describes the mean sampled forward cost, and no sampled estimate certifies optimality (Tier 3).
6 Graceful Shutdown
Section titled “6 Graceful Shutdown”A shutdown request ends the training loop at an iteration boundary, with the latest completed iteration’s policy persisted.
Guarantee: The policy at the moment of termination is usable. The loop reads a request once per iteration, after its lower bound and before its stop decision, and completes that iteration before it stops, so the cuts and bounds of every iteration that ran are recorded and no partial iteration is left behind. The graceful-shutdown guarantee is a special case of the broader provenance commitment described in Determinism & Provenance §5: the output artefacts are always in a consistent state, whether the run reached a configured stopping rule or a shutdown request ended it.
Unconditional: Graceful shutdown is not a configurable rule; it is an unconditional
safety property of the training loop. It is not listed in stopping_rules in the case
configuration and is not subject to stopping_mode combination logic. Which interfaces
can issue a request is listed in the Configure tab under
Implementation in Novomodelo.
Trade-off: A request takes effect at an iteration boundary, and the iteration that reads it runs to completion, so a request takes effect after the iteration that reads it completes, not when it is issued. The alternative (immediate termination) would leave the policy state inconsistent.
7 Combining Rules
Section titled “7 Combining Rules”Mode: "any" (default):
First rule to trigger causes termination.
Mode: "all":
The conjunction runs over the rules other than the iteration limits and is false when there are none; is the largest iteration limit, and a smaller iteration limit has no effect.
Graceful shutdown is independent of the stopping_mode combination logic — it terminates
the training loop regardless of whether any or all of the configured rules have triggered.
Implementation in Novomodelo
Section titled “Implementation in Novomodelo”The methodology above defines when each rule stops training; the tab below links the configuration entries for the rules, states which interfaces can issue a shutdown request, and lists what a stopped run records.
Novomodelo’s stopping settings are two keys of the training object of
config.json: stopping_rules, the list of rules a run can stop on, and
stopping_mode, how they combine. This tab gives one combined set and links
the configuration entries that list every field and default; each rule’s stop
condition stays in the methodology sections above
(§2 Iteration Limit,
§3 Time Limit,
§4 Bound Stalling and
§5 Gap-Based Stopping).
training.stopping_rules — Rules
Section titled “training.stopping_rules — Rules”The configuration reference lists the fields of each rule
(iteration_limit,
time_limit,
bound_stalling and
gap), and the values and default of
training.stopping_mode.
training.stopping_mode — Combination
Section titled “training.stopping_mode — Combination”The following set stops when the gap closes or at 500 iterations, whichever comes first:
{ "training": { "stopping_rules": [ { "type": "iteration_limit", "limit": 500 }, { "type": "gap", "tolerance": 1.0, "relative_tolerance": 0.1 } ], "stopping_mode": "any" }}Its gap rule needs an enumerated training forward pass; a sampled run stops
as described in Sampled runs.
Shutdown Requests
Section titled “Shutdown Requests”A shutdown request is not an entry of stopping_rules, and stopping_mode
does not apply to it. During training, novomodelo run turns a SIGTERM or SIGINT
into one, which training reads
once per iteration, after the lower bound, before the stop decision: training
stops at the end of the iteration that reads it, writes the training outputs
and the policy checkpoint, skips the configured simulation, and exits 5
(CLI Reference — Exit Codes). A signal
that reaches any rank of a multi-rank run stops every rank at the same
iteration, and a second SIGINT ends a single-process run at once. Before
training starts and while the simulation runs, both signals take their default
action. The Python API requests one when the on_iteration
callback returns a truthy value or raises
(Python API); the stop is asynchronous, so the run ends
at a later iteration boundary that is not fixed; a run ended this way
records graceful_shutdown unless the stopping rules or the iteration budget also ended it
at that iteration
(What a stopped run records).
What a stopped run records
Section titled “What a stopped run records”The convergence, bounds and iterations blocks of
training/metadata.json
record how the run stopped:
convergence.termination_reasonnames the first triggered rule, in the orderstopping_ruleslists them (under"all", among the rules other thaniteration_limit), when the stopping rules ended the run; otherwise it isiteration_limitwhen the run reached its iteration budget, and otherwisegraceful_shutdownwhen a shutdown request ended it. It iserrorafter a training that failed.convergence.achievedistruewhen the stopping rules ended the run and agaporbound_stallingrule triggered at that iteration.convergence.final_gap_percentisnullwhen the final lower bound is not positive.bounds.final_lower_boundandbounds.final_upper_boundare the bounds at the last iteration, anditerations.completedis the iteration count at which the run ended, earlier runs included.
Lower bound and gap under projected cuts. Under a PAR(p > 0) model with the
inflow lags projected out of the cuts, the lower bound and the gap carry the
caveat of
Storage-only cut projection;
the key and its default are in
stages[].state_variables.
Cross-References
Section titled “Cross-References”- Notation Conventions — Symbol definitions for bounds and statistical quantities
- SDDP Algorithm — Main iteration loop that evaluates stopping rules
- Cut Management — Cut generation and selection that affect convergence speed
- Upper Bound Evaluation — The exact (deterministic) and statistical upper-bound estimators; the exact bound is what the gap rule compares against
- Risk Measures — Risk-averse formulations that affect bound interpretation
- Determinism & Provenance — Provenance commitment that the graceful-shutdown guarantee is a special case of