Skip to content

Stopping Rules

This chapter defines the available stopping rules for the Novomodelo SDDP solver and how they combine. It covers iteration limits, time limits, bound stalling, and the recommended gap-based stopping criterion.

SDDP can terminate based on multiple criteria. Each rule is evaluated independently, and the stopping_mode determines how they combine:

  • "any": Stop when any rule triggers (OR logic)

  • "all": Stop when all rules other than the iteration limits trigger at the same iteration (AND logic); the largest iteration limit caps the run

Evaluation:

STOP  ⟺  k≥kmax\text{STOP} \iff k \geq k_{max}

where kk is the current iteration and kmaxk_{max} is the limit.

Purpose: Safety bound to prevent infinite loops.

Evaluation:

STOP  ⟺  telapsed≥tmax\text{STOP} \iff t_{elapsed} \geq t_{max}

Wall-clock time is checked at the end of each iteration.

Evaluation:

Track the lower bound z‾k\underline{z}^k over iterations. Once τ\tau bounds have been recorded, compute the relative improvement over the window of the last τ\tau recorded lower bounds, z‾k−τ+1,…,z‾k\underline{z}^{k-\tau+1}, \ldots, \underline{z}^k (the iterations parameter):

Δk=z‾k−z‾k−τ+1max⁡(1,∣z‾k∣)\Delta_k = \frac{\underline{z}^k - \underline{z}^{k-\tau+1}}{\max(1, |\underline{z}^k|)}

Stopping condition:

STOP  ⟺  ∣Δk∣<εstall\text{STOP} \iff |\Delta_k| < \varepsilon_{\text{stall}}

where εstall\varepsilon_{\text{stall}} is the bound-stalling tolerance on the relative improvement, set by the rule’s tolerance parameter.

Interpretation: The bound has plateaued — the relative improvement across the last τ\tau recorded lower bounds is below the specified tolerance, indicating diminishing returns from further iterations.

Illustrative bound evolution across iterations kk. In both panels the lower bound z‾k\underline{z}^k rises monotonically (the append-only cut pool). Left, a sampled forward pass: the upper bound zˉk\bar{z}^k is an estimate, and its confidence band reflects a sampling error that persists as the gap closes, so late in the run the band straddles the lower bound and certifies nothing. Right, an enumerated forward pass: the exact upper bound carries no band, stays above the lower bound and closes on it — the comparison the gap rule tests.

At iteration kk, the optimality gap is the difference between the upper bound zˉk\bar{z}^k and the lower bound z‾k\underline{z}^k, whichever mechanism supplies the upper bound — the statistical estimate of a sampled forward pass or the exact bound of an enumerated one (Upper Bound Evaluation):

gapk=zˉk−z‾k\text{gap}^k = \bar{z}^k - \underline{z}^k

The percent gap normalizes it by the lower bound:

100⋅gapkmax⁡(1,∣z‾k∣)100 \cdot \frac{\text{gap}^k}{\max(1, |\underline{z}^k|)}

Both bounds are in unscaled currency units — the original cost units, with the cost scaling of the stage LP undone — so gapk\text{gap}^k is in currency units and the percent gap is dimensionless. The denominator is the magnitude of the lower bound, never the upper bound, floored at one currency unit so that the ratio stays bounded when the lower bound is near zero. The per-iteration reported gap and the relative arm of the gap rule below share this denominator, so the two never disagree on what “relative gap” means.

The gap rule terminates training once the exact upper bound has closed on the lower bound.

Evaluation:

The rule evaluates the optimality gap with the exact upper bound, zˉk=zˉexact\bar{z}^k = \bar{z}_{\text{exact}}, and compares it clamped at zero, max⁡(0,  gapk)\max(0,\; \text{gap}^k); the clamp belongs to the rule’s comparison, not to the gap’s definition. In exact arithmetic the exact upper bound satisfies zˉ≥z‾\bar{z} \geq \underline{z}; the clamp only absorbs floating-point noise once the gap has closed to (numerically) zero.

Stopping condition — two arms combined by disjunction (either is sufficient):

STOP  ⟺  max⁡(0,  gapk)≤εabsor100⋅max⁡(0,  gapk)/max⁡(1,∣z‾k∣)≤εrel\text{STOP} \iff \max(0,\; \text{gap}^k) \leq \varepsilon_{\text{abs}} \quad\text{or}\quad 100 \cdot \max(0,\; \text{gap}^k) / \max(1, |\underline{z}^k|) \leq \varepsilon_{\text{rel}}

The absolute tolerance εabs\varepsilon_{\text{abs}}, in currency units, is set by the rule’s tolerance parameter. The relative tolerance εrel\varepsilon_{\text{rel}}, in percent, is set by the rule’s relative_tolerance parameter. Either tolerance may be configured alone; when both are configured, training stops as soon as either arm is satisfied.

Admissibility: the comparison above is only valid when the upper bound it uses is the exact bound, not a statistical estimate. Both of the following must hold:

  • The forward pass enumerates the scenario tree exhaustively rather than sampling it, so the upper bound is exact — carrying no sampling error — rather than a sampled approximation.
  • The risk measure is uniform across all stages: either expectation at every stage, or one CVaR measure (the same risk-aversion weight λ\lambda and tail fraction α\alpha) at every stage. Under expectation the exact bound is the probability-weighted cost ∑ℓP(ℓ) C(ℓ)\sum_{\ell} P(\ell)\, C(\ell) over every scenario ℓ\ell; under a uniform CVaR it is the nested, time-consistent risk bound over the same enumerated tree (Upper Bound Evaluation §2). A stage-varying measure — the pair (λt,αt)(\lambda_t, \alpha_t) differing across stages, where every stage with λt=0\lambda_t = 0 is the expectation whatever its αt\alpha_t — is not admissible, because no single measure then aggregates the tree.

Under a sampled forward pass, or a stage-varying measure, the exact-bound comparison the rule depends on is unavailable, and the gap rule is rejected at setup with a named validation error rather than silently evaluating against an unsound comparison.

Why recommended: unlike bound stalling (a proxy for convergence) or the iteration and time limits (safety bounds unrelated to convergence), a closed gap is a direct optimality certificate under the exact-bound regime the rule requires — an enumerated forward pass with a risk measure held uniform across stages.

The admissibility decision takes the two conditions in order: first whether the forward pass is enumerated, then whether the measure is the same at every stage. Passing both admits the gap rule, which then compares the lower bound with the exact upper bound, so that a closed gap certifies optimality under the Tier 3 hypotheses; failing either rejects the rule at setup — stop such a run with bound stalling and an iteration limit instead, as in Sampled runs.

Stopping-rule setwith a gap ruleForward passenumerated?Same measureat every stage?Gap rule admittedexact bound, certificateGap rule rejected at setupstop on bound stalling, iteration limit yesnoyesno

Under a sampled forward pass the gap rule is not available. Stop a sampled run with bound stalling (§4) and an iteration limit (§2), combined so that the first to trigger ends the run (§7). The upper bound of a sampled run is a statistical estimate: its confidence interval describes the mean sampled forward cost, and no sampled estimate certifies optimality (Tier 3).

A shutdown request ends the training loop at an iteration boundary, with the latest completed iteration’s policy persisted.

Guarantee: The policy at the moment of termination is usable. The loop reads a request once per iteration, after its lower bound and before its stop decision, and completes that iteration before it stops, so the cuts and bounds of every iteration that ran are recorded and no partial iteration is left behind. The graceful-shutdown guarantee is a special case of the broader provenance commitment described in Determinism & Provenance §5: the output artefacts are always in a consistent state, whether the run reached a configured stopping rule or a shutdown request ended it.

Unconditional: Graceful shutdown is not a configurable rule; it is an unconditional safety property of the training loop. It is not listed in stopping_rules in the case configuration and is not subject to stopping_mode combination logic. Which interfaces can issue a request is listed in the Configure tab under Implementation in Novomodelo.

Trade-off: A request takes effect at an iteration boundary, and the iteration that reads it runs to completion, so a request takes effect after the iteration that reads it completes, not when it is issued. The alternative (immediate termination) would leave the policy state inconsistent.

Mode: "any" (default):

STOP  ⟺  Rule1∨Rule2∨…\text{STOP} \iff \text{Rule}_1 \lor \text{Rule}_2 \lor \ldots

First rule to trigger causes termination.

Mode: "all":

STOP  ⟺  (Rule1∧Rule2∧…)∨k≥kmax\text{STOP} \iff (\text{Rule}_1 \land \text{Rule}_2 \land \ldots) \lor k \geq k_{max}

The conjunction runs over the rules other than the iteration limits and is false when there are none; kmaxk_{max} is the largest iteration limit, and a smaller iteration limit has no effect.

Graceful shutdown is independent of the stopping_mode combination logic — it terminates the training loop regardless of whether any or all of the configured rules have triggered.

The methodology above defines when each rule stops training; the tab below links the configuration entries for the rules, states which interfaces can issue a shutdown request, and lists what a stopped run records.

Novomodelo’s stopping settings are two keys of the training object of config.json: stopping_rules, the list of rules a run can stop on, and stopping_mode, how they combine. This tab gives one combined set and links the configuration entries that list every field and default; each rule’s stop condition stays in the methodology sections above (§2 Iteration Limit, §3 Time Limit, §4 Bound Stalling and §5 Gap-Based Stopping).

The configuration reference lists the fields of each rule (iteration_limit, time_limit, bound_stalling and gap), and the values and default of training.stopping_mode.

The following set stops when the gap closes or at 500 iterations, whichever comes first:

{
"training": {
"stopping_rules": [
{ "type": "iteration_limit", "limit": 500 },
{ "type": "gap", "tolerance": 1.0, "relative_tolerance": 0.1 }
],
"stopping_mode": "any"
}
}

Its gap rule needs an enumerated training forward pass; a sampled run stops as described in Sampled runs.

A shutdown request is not an entry of stopping_rules, and stopping_mode does not apply to it. During training, novomodelo run turns a SIGTERM or SIGINT into one, which training reads once per iteration, after the lower bound, before the stop decision: training stops at the end of the iteration that reads it, writes the training outputs and the policy checkpoint, skips the configured simulation, and exits 5 (CLI Reference — Exit Codes). A signal that reaches any rank of a multi-rank run stops every rank at the same iteration, and a second SIGINT ends a single-process run at once. Before training starts and while the simulation runs, both signals take their default action. The Python API requests one when the on_iteration callback returns a truthy value or raises (Python API); the stop is asynchronous, so the run ends at a later iteration boundary that is not fixed; a run ended this way records graceful_shutdown unless the stopping rules or the iteration budget also ended it at that iteration (What a stopped run records).

The convergence, bounds and iterations blocks of training/metadata.json record how the run stopped:

  • convergence.termination_reason names the first triggered rule, in the order stopping_rules lists them (under "all", among the rules other than iteration_limit), when the stopping rules ended the run; otherwise it is iteration_limit when the run reached its iteration budget, and otherwise graceful_shutdown when a shutdown request ended it. It is error after a training that failed.
  • convergence.achieved is true when the stopping rules ended the run and a gap or bound_stalling rule triggered at that iteration.
  • convergence.final_gap_percent is null when the final lower bound is not positive.
  • bounds.final_lower_bound and bounds.final_upper_bound are the bounds at the last iteration, and iterations.completed is the iteration count at which the run ended, earlier runs included.

Lower bound and gap under projected cuts. Under a PAR(p > 0) model with the inflow lags projected out of the cuts, the lower bound and the gap carry the caveat of Storage-only cut projection; the key and its default are in stages[].state_variables.

  • Notation Conventions — Symbol definitions for bounds and statistical quantities
  • SDDP Algorithm — Main iteration loop that evaluates stopping rules
  • Cut Management — Cut generation and selection that affect convergence speed
  • Upper Bound Evaluation — The exact (deterministic) and statistical upper-bound estimators; the exact bound is what the gap rule compares against
  • Risk Measures — Risk-averse formulations that affect bound interpretation
  • Determinism & Provenance — Provenance commitment that the graceful-shutdown guarantee is a special case of