Scenario Generation
Purpose
Section titled “Purpose”This chapter defines the Novomodelo scenario generation methodology. It covers reproducible noise generation, correlated noise generation, the opening tree concept, the sampling scheme abstraction that governs how scenarios are selected in forward and backward passes, external scenario integration, load and non-controllable source scenarios, and enumerated scenario trees. For the mathematical definition of the PAR(p) model itself, see PAR(p) Inflow Model.
1. The Inflow Model
Section titled “1. The Inflow Model”Scenario generation samples the applied PAR(p) model of PAR(p) Inflow Model, which defines the model and its estimation, the innovation-scale closure and the validation invariants. The inflow-lag seed that starts every trajectory is the lag state of §4.5.
2. Noise Generation
Section titled “2. Noise Generation”2.1 Correlated Noise Generation
Section titled “2.1 Correlated Noise Generation”Hydros within the same correlation group share spatially correlated noise. Each group carries its own correlation matrix.
The generation process:
- Independent sampling — Generate independent standard normal samples, one per entity in the correlation group
- Spectral transform — Pre-decompose the correlation matrix via eigendecomposition during preprocessing. The spectral factor transforms independent noise into the correlated PAR innovation: . Negative eigenvalues are clipped to zero, which gives the nearest positive-semidefinite matrix in the Frobenius norm; its diagonal can exceed one (see PAR(p) Inflow Model). Spectral factorisation is the only method Novomodelo implements, and the same section gives the rationale.
- Entity assignment — Each entity in the group receives its own correlated noise value from the transformed vector.
Entities in different correlation groups are independent of each other. Entities not assigned to any group receive independent noise.
For the spectral factorisation rationale and the full eigendecomposition derivation, see PAR(p) Inflow Model.
2.2 Reproducible Sampling
Section titled “2.2 Reproducible Sampling”Noise generation must produce identical results regardless of MPI rank assignment, thread scheduling, or restart. This is achieved through deterministic seed derivation:
- A base seed is fixed for the study
- Each draw maps to a unique derived seed via a deterministic hash function of a tuple that identifies it. The tuple depends on the sampling method and on the pass: for example (iteration, scenario, stage) for the in-sample forward choice of an opening and (opening, stage) for a Monte Carlo opening-tree draw; Determinism & Provenance §3 lists the tuple of every pass
- The RNG state is initialized from this derived seed before generating the noise vector for that tuple
This design ensures:
- Cross-rank reproducibility — The same scenario/stage produces the same noise regardless of which MPI rank processes it
- Restart reproducibility — Resuming from a checkpoint produces identical noise for subsequent iterations
- Order independence — Results are identical regardless of the order in which scenarios or stages are processed
Each seed is a hash of a fixed-width tuple of integers, encoded little-endian so that the result is identical on every platform. The Implementation notes tab under Implementation in Novomodelo gives the tuple layouts.
2.3 Opening Tree
Section titled “2.3 Opening Tree”The backward pass in SDDP evaluates an aggregated cut at a node by solving all of that node’s openings — the within-node opening set defined in Policy Graphs section 2. These openings must be identical across all iterations — the backward pass always “sees the same tree” at every node. Novomodelo fixes each node’s opening set before training and visits all of it in every backward pass. This fixed, per-node structure is the opening tree. On the implicit stage chain — one (unnamed) node per stage, the case the rest of this chapter documents — collapses to the per-stage branching factor used below.
The tree is generated once before training begins and remains fixed throughout. Generation produces noise vectors for each node — per stage on the implicit chain; the backward pass iterates over them at every iteration of the algorithm.
Each node’s opening count is configured per study — the per-stage branching factor on the implicit chain. Uniform branching ( for all ) is the common case in standard SDDP, but per-node variable counts are supported. The backward pass solves one subproblem per opening at each node it visits, so its cost grows in proportion to ; every count yields valid cuts, so the choice affects the cost of an iteration and the speed of convergence, not the validity of the lower bound.
Tree generation:
- Before the first SDDP iteration, generate noise vectors per stage , producing a fixed opening tree with total element count
- Each noise vector is generated from the base seed using deterministic seed derivation, per (opening, stage) under Monte Carlo sampling and per stage under a batch method (Latin hypercube or quasi-Monte Carlo), which draws all of a stage’s openings together — ensuring reproducibility across restarts and MPI configurations (see section 2.2)
- Correlation is applied per the active profile for each stage (see section 2.4)
- In the opening tree, consecutive stages sharing a noise group — keyed by season and year, as when several weekly stages fall inside one monthly season — receive one shared noise draw per opening: the group’s first stage draws its noise as in step 2 and each later stage in the group copies that stage’s already-correlated noise. A stage without a season gets its own group; in a uniform monthly study every group is unique and every stage draws independently. See Multi-Resolution Studies for the mechanism’s role in mixed-resolution horizons.
Each node’s opening set is drawn from the tree generated for its stage — on the implicit chain this is the entire stage- set of vectors produced above.
Backward pass usage: At each node , the backward pass iterates over all of — noise vectors at stage on the implicit chain — solving one subproblem per opening, then aggregates the resulting cuts, composed across successor nodes per the between-node edge weights (see Policy Graphs section 2). Because the tree is fixed, every iteration produces cuts that refine the same set of future cost scenarios.
Forward pass usage (InSample only): When the InSample sampling scheme is
active (see section 3.2), the forward pass samples a random opening
at each node it visits — a random index
on the implicit chain — and uses its
corresponding noise vector from the opening tree. The other
sampling schemes do not access the opening tree: OutOfSample draws fresh
noise with each stage’s sampling method, plain Monte Carlo at selective and
historical-residual stages (sections 2.3a and 3.2), and External and
Historical use entirely separate data sources; the forward pass noise path is
governed by the sampling scheme abstraction (section 3).
The opening tree uses stage-major ordering so that all of a stage’s noise vectors — the set its node(s) draw from — are contiguous in memory, enabling linear access during the backward pass. The noise vector for a given (stage, opening) pair is read directly without iteration over the tree structure.
The same tree is generated bit-identically on every MPI rank because both the seed derivation of step 2 and the noise-group assignment are pure functions of globally known inputs — the base seed and the stage calendar of the broadcast system. For the determinism guarantees that follow from this property, see Determinism & Provenance.
2.3a Sampling Method and Opening Tree Generation
Section titled “2.3a Sampling Method and Opening Tree Generation”Each stage selects the sampling method that generates its openings — the set each stage’s node(s) draw their from. The method is orthogonal to the forward sampling scheme of section 3: the method governs how the opening tree is populated with noise vectors; the scheme governs which noise source the forward pass uses. The methods are:
- Plain Monte Carlo — the default. Each opening’s noise vector has independent standard normal components, the openings are equiprobable, , and the correlation factor of section 2.1 then turns each opening into its correlated noise vector .
- Latin hypercube — For each coordinate, the unit interval is split into equal strata and one uniform point is drawn in each stratum; an independent random permutation per coordinate assigns the strata to the openings, and maps each point to a standard-normal value.
- Scrambled Sobol — The openings are points of a Sobol low-discrepancy sequence with random linear scrambling of each coordinate, mapped by . The sequence is bounded in dimension: a stage whose noise vector has more coordinates than the sequence supports is rejected.
- Scrambled Halton — The openings are points of a Halton sequence, each coordinate on its own prime base, with random digit scrambling, mapped by .
- Selective — The openings are supplied as a pre-built tree, which provides the openings of every stage, whatever method the stage selects. Novomodelo generates none, and a stage selecting this method without a supplied tree is rejected.
- Historical residuals — Each opening carries the standardized inflow noise of one window of the historical record, which already embeds the record’s cross-hydro correlation, so no correlation factor is applied (see §2.5).
In the first four methods the correlation factor of section 2.1 is applied after the draws, so the stratification and the low-discrepancy properties hold for the uncorrelated draws; denotes the inverse of the standard normal distribution function. Under every method a stage’s openings are equiprobable; where section 2.5 or 4.2 caps a stage’s opening count below , the capped count takes the place of in and in the number of strata.
Per-stage variation. The method can vary per stage, enabling mixed strategies in a single run.
2.4 Time-Varying Correlation Profiles
Section titled “2.4 Time-Varying Correlation Profiles”The correlation structure can vary across stages via the profile + schedule pattern: the correlation profiles and their stage schedule.
- Profiles — Named correlation configurations (for example, a default profile and profiles for the wet and dry seasons), each defining correlation groups and matrices
- Schedule — An optional assignment of stages to profile names. A stage the schedule does not list uses the default profile.
During preprocessing, the spectral decomposition is computed once per profile (not per stage). At runtime, the scenario generator looks up the active profile for the current stage via the schedule and uses its pre-computed spectral factor.
2.5 Historical-Residual Openings
Section titled “2.5 Historical-Residual Openings”Under any forward scheme, a generated opening tree fills a stage that uses the historical-residual method from the window pool and the applied PAR model of Historical Window Pool, with the same derived seed rooting each window’s lag chain. The pool is the one built from the candidate years of training, and it takes the place of the noise generation of section 2.1 at that stage. A supplied tree provides the openings of every stage, so this section does not apply to it (section 2.3a).
Opening count. The branching factor is capped by the pool size and, when a class is external in training, by the stage’s number of external scenarios (§4.2). Without an external class the opening count is
and with one, is the smaller of this minimum and the stage’s number of external scenarios. The study is warned when the pool is smaller than .
Draw. Opening at stage reads window , where is the opening-tree seed derivation of section 2.2 from the base seed, the opening index and the stage. The draw is independent per stage and with replacement: opening can read different windows at different stages, and two openings of one stage can read the same window. In the opening tree, a later stage of a noise group copies its group’s first stage instead of drawing (section 2.3). Each opening carries probability , whichever window it drew.
Content. The opening that reads window carries that window’s standardized inflow noise for every hydro . All hydros’ residuals come from the same window, so they already carry the record’s cross-hydro correlation, and no spectral factor is applied. The load and NCS parts of the opening are zero, so at these stages load and NCS take their zero-noise realizations: the mean, truncated at zero for load and clamped to for NCS, as in sections 5.1 and 5.4.
Application. An opening is a residual, not an inflow. The backward pass applies it through the stage LP at the trial point’s lag state, like any other opening (section 2.3): the residual enters the PAR equation with the trial point’s lags, not with the window’s own. §4.5 states when the two give the same inflow.
3. Sampling Scheme Abstraction
Section titled “3. Sampling Scheme Abstraction”3.1 Three Orthogonal Concerns
Section titled “3.1 Three Orthogonal Concerns”The SDDP algorithm has three concerns that govern how scenarios are handled during training. Novomodelo formalizes these as distinct abstractions, following the design established by SDDP.jl (Dowson & Kapelevich, 2021):
| Concern | Abstraction | What It Controls | Default |
|---|---|---|---|
| Which noise is selected at each forward pass stage | Sampling Scheme | Forward scenario selection | InSample |
| Which LP model is solved in the forward pass | Forward Pass Model | Training LP vs alternative model | Default |
| Which noise terms are evaluated in the backward pass | Backward Sampling | Branching completeness | Complete |
These three concerns are orthogonal: the choice for one does not constrain the choice for another. Two of them are fixed — the forward pass model is Default and backward sampling is Complete — so the sampling scheme is the one a study chooses.
Forward and backward noise source separation is a natural consequence of this design: the sampling scheme controls the forward pass noise source, while backward sampling always draws from the fixed opening tree (section 2.3). These two sources may differ — for example, the forward pass may sample from external scenarios while the backward pass evaluates all openings through the applied PAR model (§4.2).
Under the External and Historical schemes the forward trajectories follow the external or historical law, while the cuts and the lower bound are built on the opening tree, whose openings the stage LP turns into realizations with the applied model (§4.2). The forward-cost estimate is then the policy’s expected cost under the forward law, and the lower bound bounds the value of the tree problem from below, so the gap of Stopping Rules compares two problems and need not close or stay non-negative.
3.2 Forward Sampling Schemes
Section titled “3.2 Forward Sampling Schemes”The sampling scheme determines how the forward pass selects a scenario realization at each stage. Novomodelo supports four sampling schemes, chosen independently for each stochastic class (inflow, load, NCS); the Historical scheme applies to the inflow class only:
InSample (Default)
Section titled “InSample (Default)”At each stage , sample a random index and use the corresponding noise vector from the fixed opening tree (section 2.3). The PAR model dynamics equation is embedded in the LP as a constraint; the solver evaluates the inflow realization implicitly when it solves the LP with the fixed noise.
- Noise source: Opening tree (same as backward pass)
- Use case: Standard SDDP training — forward and backward passes see the same noise distribution
This is SDDP.jl’s InSampleMonteCarlo: the forward pass samples from the
same noise terms defined in the model.
OutOfSample
Section titled “OutOfSample”The forward pass draws from independently generated Monte Carlo noise that is distinct from the opening tree noise. Each iteration and trajectory draws fresh noise from the applied stochastic model of each class that selects out-of-sample, from a seed independent of the opening tree’s. The draws use their stage’s sampling method (section 2.3a), spread over the iteration’s trajectories: the Latin-hypercube strata and the scrambled sequence points are shared out among the trajectories as the tree shares them out among the openings. At selective and historical-residual stages the draws are plain Monte Carlo. The correlation factor of section 2.1 is then applied to every draw. The backward pass uses the same fixed opening tree as InSample.
- Noise source: Independently generated Monte Carlo noise (not from the opening tree)
- Backward pass interaction: The backward pass uses the fixed opening tree, whose openings the stage LP turns into realizations with the same applied model
- Use case: Out-of-sample forward evaluation to reduce in-sample bias
This corresponds to SDDP.jl’s OutOfSampleMonteCarlo.
External
Section titled “External”The forward pass draws from user-provided per-class scenario data (for example, a set of external inflow scenarios). Each forward trajectory replays one external scenario for all its stages, chosen with replacement by a hash of the training iteration and trajectory indices from its own fixed base. A policy-graph node whose openings point into its stage’s external realization column uses that column instead (see Policy Graphs).
- Noise source: External scenario values (not the opening tree)
- Realization computation: External values must always be inverted to noise terms (epsilon) before they can be used in the LP (see section 4.3).
- Backward pass interaction: The backward pass still uses the fixed opening tree, whose openings the stage LP turns into realizations with the applied model: the PAR model for inflow, and a mean and standard deviation per entity for load and NCS. Under the External scheme that model takes the external samples’ moments for the entities of §4.2.
- Use case: Training with imported Monte Carlo scenarios or stress-test scenarios
Historical
Section titled “Historical”Each forward trajectory replays one window of the historical window pool for all its stages. The window of a trajectory is , a hash of the training iteration and trajectory indices from a fixed base, so windows are drawn with replacement and can repeat within and across iterations.
- Noise source: The standardized noise of the window pool, one window per trajectory
- Realization computation: Historical values must be inverted to noise terms (epsilon) before use in the LP. The same noise inversion procedure applies (see section 4.3). The lag chain seeding the inversion is rooted at the derived inflow-lag seed, the lag state of , and advanced by the per-stage transitions, not at the year-preceding raw historical lags. The reason and consequences are spelled out in §4.5.
- Backward pass interaction: The backward pass uses the fixed opening tree, whose openings the stage LP turns into inflows with the study’s PAR model as its inputs define it; the Historical scheme substitutes nothing.
- Use case: Policy validation against observed conditions, historical replay analysis
This matches the semantics of SDDP.jl’s Historical sampling scheme.
Historical Window Pool
Section titled “Historical Window Pool”The Historical scheme draws its replayed sequences from a pool of windows of the historical record. A window is a year , the year of stage 1’s season occurrence (for a weekly cycle, its ISO week-numbering year). For stage , window reads the occurrence of season in year , with the offset of stage ‘s season-occurrence year from stage 1’s, so .
Window pool. A year is admissible when the record holds an observation of every hydro for every occurrence the window covers:
with the record’s observation of hydro for the occurrence of season in year . The candidate years are the years of the record, or a list of years the study supplies; a candidate outside is dropped, and an empty refuses the study (§4.3 lists the refusals). The replay’s lag chain starts from the derived inflow-lag seed, the lag state of (see §4.5).
Inversion. Window ‘s noise at stage inverts the PAR model on its observation:
where is window ‘s lag chain: the derived seed before stage 1, then window ‘s own observations (§4.3). When , the noise is if the observation equals the deterministic value to within a numerical tolerance; otherwise the inversion has no finite value and the study is refused.
For every hydro, each window must cover the season occurrence of every study stage.
3.3 Forward Pass Model
Section titled “3.3 Forward Pass Model”The forward pass model determines which LP is solved at each stage during the forward pass. Novomodelo implements one model:
- Default — Solve the training LP (the convex LP used for cut generation) with the scenario realization from the sampling scheme. This is the standard SDDP forward pass.
3.4 Backward Sampling
Section titled “3.4 Backward Sampling”The backward sampling scheme determines which noise terms are evaluated at each stage during the backward pass. Novomodelo implements one scheme:
- Complete — Evaluate all noise vectors from the fixed opening tree. This is the standard SDDP backward pass that guarantees proper cut generation by considering every branching.
Every backward pass evaluates the full opening set; a Monte Carlo backward sampling variant would instead sample openings with replacement.
4. External Scenario Integration
Section titled “4. External Scenario Integration”4.1 External Scenario Sources
Section titled “4.1 External Scenario Sources”Novomodelo supports external scenarios — pre-generated realizations indexed by scenario, of which each forward trajectory replays one (§3.2) — as an alternative forward-pass noise source. External scenarios can drive the forward pass in both training and simulation. For each class that is external in training, the external samples’ own mean and standard deviation replace the stored statistics of that class’s entities: every load bus with load statistics or external values for the load class, every NCS source with availability statistics or external values for the NCS class, and, for the inflow class, every hydro whose inflow model has AR order 0 and no annual component. This includes the degenerate deterministic case, a constant column with standard deviation zero (see §4.4). A hydro with AR order above 0 or an annual component keeps the mean, standard deviation and coefficients of the study’s PAR model.
| Use Case | Description |
|---|---|
| Historical replay | Use actual historical inflows |
| Monte Carlo import | Pre-generated scenarios from external tool |
| Stress testing | Specific drought/flood scenarios |
4.2 Backward Pass with External Forward Scenarios
Section titled “4.2 Backward Pass with External Forward Scenarios”When the External or Historical sampling scheme is active during training,
the forward and backward passes use different noise sources. The backward
pass still requires proper probabilistic branchings for valid cut generation,
so it evaluates the fixed opening tree (section 2.3) through the applied model.
The lifecycle is:
- Moment substitution — For each class under the External scheme, the mean and the population standard deviation of that class’s external values, per entity and stage, replace the stored statistics of the class’s entities: every load bus with load statistics or external values for the load class, every NCS source with availability statistics or external values for the NCS class, and, for the inflow class, every hydro whose inflow model has AR order 0 and no annual component; the divisor is the stage’s number of external scenarios. Nothing is fitted to the external values: a hydro with AR order above 0 or an annual component keeps the study’s PAR model unchanged, AR structure included. The forward-pass standardization uses the same moments, so the two never disagree about whether a series is deterministic. The Historical scheme substitutes nothing.
- Opening tree generation — Generate the fixed opening tree (section 2.3). Its openings are standardized noise, which the stage LP turns into realizations with the applied model — the study’s model after step 1 — so the backward branchings carry the external samples’ moments for the entities of step 1.
- Training — Forward pass samples from external data; backward pass evaluates all openings through the applied model.
When a class is external, a stage’s generated opening count is at most that stage’s number of external scenarios; at a historical-residual stage the window-pool bound of §2.5 applies as well, and the opening count is the smallest of , and the stage’s number of external scenarios.
Rationale: The backward pass needs proper probabilistic branchings for valid cuts, so it samples the applied model through the opening tree rather than the external scenarios themselves; for the entities of step 1, the substituted moments carry the external samples’ mean and spread into those branchings.
4.3 Noise Inversion for External and Historical Scenarios
Section titled “4.3 Noise Inversion for External and Historical Scenarios”When using external or historical scenarios, Novomodelo must compute the implied noise values that would have generated those inflows under the PAR model. This is required because the backward pass needs noise values to construct the appropriate RHS perturbations for cut generation, and the lower bound and upper bound must be evaluated at the same for the gap to be meaningful.
Given target inflow at stage for hydro with season :
where is the precomputed base value.
All quantities (, , ) are in their respective units as described in PAR(p) Inflow Model. The AR coefficients are in original units; is the derived residual std. Under PAR(p)-A, the upper index of the lag sum equals the full width of the AR-dynamics row (twelve lags, or the classical order if larger, when any hydro carries the annual component), not the classical AR order — the annual coefficient is spread across the trailing lag positions and must enter the deterministic component of the inversion to keep the implied noise consistent with the LP’s AR row.
The inversion proceeds sequentially through stages (each stage updates the lag buffer for the next):
- Initialise the lag buffer from the derived inflow-lag seed — the observed inflows before the study start, cast from the windowed inflow history and shadowed day-wise by the initial conditions’ recent observations wherever the two overlap (see §4.5 for the reason this is the only admissible seed).
- For each stage : compute the deterministic PAR component, solve for , update the lag buffer with .
- Validate the inverted noise:
- When , the inverted noise is if equals the deterministic component to within a numerical tolerance, and otherwise has no finite value, which refuses the study (the PAR model says the series is deterministic, but the scenario disagrees).
- Building the inverted window pool, which serves both the historical replay and the historical-residual openings (§2.5), refuses the study when a study stage has no season, when an inverted noise value is the mismatch above or is not a number, or when the window pool is empty. Error Codes lists the messages and the exit codes.
4.4 Deterministic (σ = 0) External Columns
Section titled “4.4 Deterministic (σ = 0) External Columns”A class’s external column may be constant — every realization identical, standard deviation zero. Because the class statistics are derived from the samples themselves (§4.1), Novomodelo reads that as a genuinely deterministic series rather than a malformed one, and accepts it for load, for non-controllable sources, and for an inflow with an order-0 model (no autoregressive lag and no annual component): the deterministic base is exactly that constant value.
A constant column is rejected only for an inflow whose model has an autoregressive lag or an annual component. A single deterministic value cannot stand in for such a model: it would have to equal the model’s own stage-by-stage deterministic PAR output — which depends on the evolving lag state and is not reconstructed from a flat column — so the load fails, naming the class, stage, and entity. This is distinct from the §4.3 inversion anchor check: it is a load-time admissibility rule on the column, not a per-stage residual test. (The concrete rejection message is a software-layer concern — see Error Codes.)
4.5 x₀ Consistency Under Historical Replay
Section titled “4.5 x₀ Consistency Under Historical Replay”The historical and external schemes share the same SDDP forward pass: the sampler returns a standardised noise residual and the LP reconstructs the realised inflow from together with the lag state carried in the state vector. The lag state at stage 0 is taken uniformly from the derived inflow-lag seed for every scenario, so all forward replays share a single hydrological starting tendency.
For the implied on a historical window to reconstruct the raw historical observation exactly when the LP starts from the same , the inversion lag chain must also be rooted at that same derived seed — not at the year-preceding raw historical inflows of the window being replayed. If the inversion is seeded from the window-preceding lags while the LP starts from the derived seed, the two paths build their lag chains from different roots and produce a systematic per-stage offset of
that propagates through every stage. This offset prevents exact replay of the historical observation even at stage 0 and shows up as a structural, typically-negative gap between the forward upper bound and the lower bound that does not close with iteration count.
To eliminate this gap, the inversion runs as a rolling lag chain seeded from the derived seed and advanced at every stage with the same per-stage lag transitions the LP uses (see Multi-Resolution Studies), for monthly, sub-monthly and multi-monthly stages alike.
The cross-scheme implications are:
- Historical scheme. Every scenario starts the forward LP from the derived seed; the implied rebuilds the raw historical observation to within floating-point precision at every stage; the LB and UB evaluate at the same . The reconstruction is exact under the inflow non-negativity methods that pass the noise through. Under the truncation methods, a hydro’s first negative observation replays as zero inflow, and from that stage on the hydro’s replayed lag chain departs from the observed one (see Inflow Non-Negativity Solution Methods).
- Historical-residual openings. When training replays history, the openings of a historical-residual stage (§2.5) come from the same window pool, applied PAR model and derived seed as the replay. At a trial point whose lags are window ‘s chain , an opening that drew window at the stage reproduces that window’s observation there (under the truncation methods, a negative observation replays as zero, as above); at any other lag state the same residual enters the PAR equation with that state’s lags and in general yields a different inflow.
- External scheme. Same rolling-chain machinery for inflow — the supplied target series is inverted relative to the derived seed, preserving LB/UB consistency at across all forward replays; load and NCS carry no lag structure, and their values are standardized by the class’s applied mean and standard deviation alone. Under the truncation methods, a hydro’s first negative target value replays as zero inflow, and from that stage on the hydro’s replayed lag chain departs from the supplied series (see Inflow Non-Negativity Solution Methods).
- InSample / OutOfSample. Unaffected — their forward pass replays no observed series, so it has no reconstruction to keep consistent. An OutOfSample pass draws fresh noise (§3.2), plain Monte Carlo even at a historical-residual stage. An InSample pass samples the opening tree, so at a historical-residual stage it reads an inverted window residual (§2.5) and applies it at the trajectory’s own lag state, as the backward pass does.
4.6 External Scenarios in Simulation
Section titled “4.6 External Scenarios in Simulation”Simulation selects a class’s external scenario or historical window with the same hashes as training (§3.2), with the iteration argument held at a fixed value that no training iteration takes, so a simulated scenario’s draw depends on its scenario index only. A policy-graph node that points into its stage’s external realization column uses that column.
5. Load and Non-Controllable Source Scenarios
Section titled “5. Load and Non-Controllable Source Scenarios”5.1 Load Uncertainty Model
Section titled “5.1 Load Uncertainty Model”When load statistics are provided, the system generates stochastic load realizations per (bus, stage) using the stored mean and standard deviation. Load has no autoregressive structure: a load realization is the stored mean plus the standard deviation times a standard-normal noise, truncated at zero:
where is the bus’s load noise (§5.3). When the truncation raises the mean of above , the more so the smaller . Unless the load class is external in training, a bus whose standard deviation is zero at every stage carries itself, without the truncation.
Without load statistics, every bus carries zero load, unless the load class is external in training: the buses of its external scenarios then take the external samples’ moments (§4.1).
5.2 Block Load Factors
Section titled “5.2 Block Load Factors”The base load realization is a stage-level value in MW. Block-level loads are obtained by applying multiplicative block factors:
where is the block factor of bus at stage and block . Where no block factor is given, a block factor of one applies.
5.3 Load Correlation
Section titled “5.3 Load Correlation”A correlation group holds entities of one class only: hydros, load buses or non-controllable sources. Load noise is correlated with the other load buses of its group through the factorisation of section 2.1, and is independent of inflow and NCS noise.
5.4 Non-Controllable Source Availability
Section titled “5.4 Non-Controllable Source Availability”The availability of non-controllable source at a stage is the ratio
where and are the mean and standard deviation of the unclamped availability factor at the stage and is its noise. The ratio is a dimensionless fraction of the installed capacity: the source’s generation cap in a block is the installed capacity times times the block factor (Equipment-Specific Formulations §6 states that bound and curtailment). With the clamp raises the mean of above when and lowers it when . NCS noise is correlated only with the other sources of its group (§5.3).
The model applies to a source with a stochastic availability model; a source without one takes its stage’s available generation, the installed capacity unless the stage sets it, in place of the installed capacity times , as Equipment Formulations §6 states.
6. Enumerated Scenario Trees
Section titled “6. Enumerated Scenario Trees”Sampling evaluates a node’s opening set by drawing from it. An enumerated pass instead visits every node and every root-to-leaf path of a finite policy graph once, with no sampling, so its bounds are those of the declared tree: the upper bound is computed over every path rather than estimated from a sample (Upper Bound Evaluation §2). Its structural requirements — a singleton opening set at every node, and in-degree 1 everywhere, so the graph is a pure tree — are those of Policy Graphs §6. Branching is therefore declared as sibling nodes, each with one opening, and the deterministic trunk with branching only at the final stage (Maceira et al., 2002) is a chain of single-successor nodes ending in a terminal fan of sibling nodes.
With a common branching count per stage, the number of root-to-leaf paths is the product of the per-stage branching counts, so it grows exponentially with the number of branching stages. Enumeration suits short horizons and trees that branch at few stages.
Implementation in Novomodelo
Section titled “Implementation in Novomodelo”The methodology above defines how scenarios are generated and sampled; the tabs below cover how Novomodelo’s software surface configures the schemes, seeds and inputs, and how it implements seed derivation, the inversion lag chain and its warnings.
You set the sampling scheme, the opening-tree method and the seeds in config.json and stages.json, and supply the input files named below. This tab maps the concepts of the methodology above to those keys and files; Configuration and Case Format own the field tables.
Sampling schemes
Section titled “Sampling schemes”§3.2 defines the four forward sampling schemes. Each stochastic class (inflow, load, ncs) takes its own scheme value in training.scenario_source, and a class without one is in_sample.
| Scheme | Per-class scheme value | Required inputs |
|---|---|---|
| InSample | in_sample | Uncertainty models or inflow history |
| OutOfSample | out_of_sample | The same inputs as in_sample |
| External | external | The class’s external scenario file: scenarios/external_inflow_scenarios.parquet, scenarios/external_load_scenarios.parquet or scenarios/external_ncs_scenarios.parquet |
| Historical | historical (inflow only) | scenarios/inflow_history.parquet and season_definitions in stages.json |
For out_of_sample, the forward draw is fresh noise from a seed independent of the tree’s, drawn with the stage’s sampling_method (plain Monte Carlo at selective and historical_residuals stages), with correlation applied to every draw.
historical is valid for the inflow class only. The key training.scenario_source.historical_years restricts the pool of historical windows; see scenario_source. simulation.scenario_source sets the simulation schemes in the same way, and when it is absent simulation uses training.scenario_source. out_of_sample and external need the section’s scenario_source.seed, and historical needs none; simulation.scenario_source.seed seeds the simulation’s out-of-sample draws (see Seeds).
Opening-tree methods
Section titled “Opening-tree methods”§2.3a defines the methods. Each study stage selects one with sampling_method in stages.json.
| Method | stages[].sampling_method | Requires |
|---|---|---|
| Plain Monte Carlo | saa (the default) | Nothing further |
| Latin hypercube | lhs | Nothing further |
| Scrambled Sobol | qmc_sobol | A noise dimension within the Sobol limit (see the Implementation notes tab) |
| Scrambled Halton | qmc_halton | Nothing further |
| Selective | selective | A supplied opening tree: training.scenario_source.openings set to {"source": "file"}, which reads scenarios/noise_openings.parquet |
| Historical residuals | historical_residuals | scenarios/inflow_history.parquet and a season on every study stage, when the opening tree is generated |
| Key | Seeds | Default |
|---|---|---|
training.tree_seed | The opening tree and the in-sample opening draws | 42 when absent; a negative value is used by its absolute value |
training.scenario_source.seed | The out-of-sample draws of training, and of simulation when simulation.scenario_source is absent | None; required when a class is out_of_sample or external |
simulation.scenario_source.seed | The simulation’s out-of-sample draws | None; required when a class of that section is out_of_sample or external |
Historical-window and external-scenario selection use no seed. Seed resolution owns the rules.
Inputs
Section titled “Inputs”scenarios/correlation.jsonholds the correlation profiles and an optionalschedulethat mapsstage_idvalues to profile names; itsmethodaccepts only"spectral". A stage the schedule does not list uses the profile nameddefault, or the only profile when the file holds one. The full schema and an example are in the Configure tab of PAR(p) Inflow Model.scenarios/inflow_seasonal_stats.parquetandscenarios/inflow_ar_coefficients.parquethold the PAR parameters, described under PAR(p) Inflow Model.scenarios/load_seasonal_stats.parquetholds the load mean and standard deviation of §5.1, andscenarios/load_factors.jsonthe block factors of §5.2; where a block factor is absent it is one.scenarios/non_controllable_stats.parquetholds the availability moments of §5.4.scenarios/external_inflow_scenarios.parquet,scenarios/external_load_scenarios.parquetandscenarios/external_ncs_scenarios.parquethold the external scenarios of §4.1, one file per class with class-specific entity id and value columns.scenarios/inflow_history.parquetholds the inflow record, andseason_definitionsinstages.jsonmaps season ids to calendar periods.recent_observationsininitial_conditions.jsonshadows the history day by day where both cover a date; the result is the derived inflow-lag seed of §4.5.scenarios/noise_openings.parquetholds a supplied opening tree, read whentraining.scenario_source.openingsis{"source": "file"}.
This tab records how Novomodelo implements §2.2 Reproducible Sampling, §2.3 Opening Tree, §2.5 Historical-Residual Openings, §4.3 Noise Inversion for External and Historical Scenarios and §4.5 x₀ Consistency Under Historical Replay, and lists the byte layouts, the lag chain, the tolerance, the warnings and the limit that the methodology leaves out.
Seed derivation
Section titled “Seed derivation”Every seed is a SipHash-1-3 digest of a little-endian byte concatenation, and the 64-bit digest seeds the PCG64 generator that draws the noise. The layouts, with the seed the first field of each:
| Seed | Bytes, in order | Length | Seed value | Draw |
|---|---|---|---|---|
| Forward | seed u64, iteration u32, scenario u32, stage u32 | 20 | training.tree_seed for the in-sample choice; a fixed base for window and external selection | The in-sample choice of an opening; with a fixed base, the choice of a historical window or an external scenario. |
| Stage batch | seed u64, stage u32 | 12 | training.tree_seed | A Latin-hypercube, Sobol or Halton opening-tree batch. |
| Opening | seed u64, opening index u32, stage u32 | 16 | training.tree_seed | A Monte Carlo opening-tree draw, and the window that a historical-residual opening reads. |
| Grouped forward | tag byte 0x01, seed u64, iteration u32, scenario u32, noise-group id u32 | 21 | The class’s forward seed | The out-of-sample forward draw of the Monte Carlo method, to which selective and historical_residuals stages fall back. |
| Class forward | tag byte 0x02, seed u64, class name bytes | 9 + name length | The phase’s forward seed | The forward seed of the load class and of the NCS class. |
The Grouped forward layout takes the class’s forward seed as its seed: the phase’s forward seed for the inflow class, and the Class forward digest of it for the load class and the NCS class. The phase’s forward seed is training.scenario_source.seed in training, and in simulation simulation.scenario_source.seed, or training.scenario_source.seed when simulation.scenario_source is absent. The Class forward layout hashes the phase’s forward seed with the class name bytes load or ncs; the inflow class has no digest of its own. The three classes therefore draw independent noise, because their seeds differ.
Window and external-scenario selection hash (iteration, scenario) with the Forward layout, stage 0, and a fixed base in place of the seed: 0x6869_7374_6f72_6963 (the ASCII text historic, most significant byte first) selects a historical window, and 0x6578_7465_726e_616c (external) selects an external scenario. No user seed enters either selection.
Inversion lag chain
Section titled “Inversion lag chain”The inversion of §4.3 advances its lag chain with the per-stage lag transitions that the LP uses. resolve_stage_lag_transitions derives the downstream PAR order and the per-stage lag transitions (StageLagTransition), and it is, in the source’s words, “The sole owner of both derivations”; the LP setup and the historical scenario library each call it once, over their own PAR model.
The library carries a fingerprint of the inflow-lag seed that rooted its inversion chain: a SipHash-1-3 digest of the seed’s lag values as little-endian f64 bytes. Novomodelo reports it as historical_library_seed_digest in training/model_provenance.json (see Metadata Files) and omits it when no library was built. A digest that differs from the one of the freshly derived seed marks a library inverted against a different x₀, whose replay is not exact.
Deterministic tolerance
Section titled “Deterministic tolerance”The σ = 0 rule of §4.3 compares the standard deviation with 0.0 exactly. At σ = 0 the inversion returns noise 0 when the target differs from the deterministic value by less than 1e-10 in magnitude, and NEG_INFINITY otherwise. The library validation refuses a NEG_INFINITY or NaN noise value. Check V2.3 covers the historical window pool and check V3.7 an external library; the error messages carry these identifiers (see Error Codes).
Warnings
Section titled “Warnings”Novomodelo logs these six warnings and continues. Each text below is fixed apart from the values at its placeholders.
-
Clamp by external scenarios. A class on the external scheme has fewer raw external scenarios at a stage than the stage’s
branching_factor, and no tighter clamp applies:External scenarios: {} raw scenarios < branching_factor {}; opening tree clamped to {} openings -
Clamp by the window pool. A
historical_residualsstage has a window pool smaller than itsbranching_factor:HistoricalResiduals: {} historical windows < branching_factor {}; opening tree clamped to {} openings -
Out-of-sample fallback. A class on the
out_of_samplescheme has stages that selectselectiveorhistorical_residuals, and the forward pass draws plain Monte Carlo noise at those stages:class '{class_str}' has {count} out-of-sample forward stage(s) selecting {methods} noise, not implemented in the forward pass; falling back to the sample-average method for stage id(s) {stage_ids:?}The
{methods}placeholder readsselective,historical_residualsorselective or historical_residuals, and the sample-average method is the plain Monte Carlo method (saa). -
Few historical windows. The historical window pool holds fewer windows than there are forward passes:
fewer windows ({}) than forward passes ({forward_passes}): historical sampling will repeat windows across forward passes -
Few windows in the historical library. Novomodelo checks the historical window pool again when it validates the library built from it, so the same shortfall logs a second warning:
historical library has fewer windows ({}) than forward passes ({}); windows will be reused across forward passes -
Few scenarios in the external library. No stage of an external scenario library holds as many scenarios as there are forward passes; each class on the external scheme is checked on its own:
external {class} library has fewer scenarios ({n_scenarios}) than forward passes ({forward_passes}); scenarios will be reused across forward passesThe
{class}placeholder readsinflow,loadorncs.
Sobol dimension limit
Section titled “Sobol dimension limit”The Sobol direction numbers cover SOBOL_MAX_DIM: usize = 21201 dimensions. A stage on qmc_sobol whose noise vector has more coordinates is rejected with the error below, where {method} is sobol.
noise dimension {dim} exceeds maximum supported dimension {max_dim} for {method}
Cross-References
Section titled “Cross-References”- PAR(p) Inflow Model — Mathematical definition, parameter set, stored vs. computed quantities, fitting procedure, spectral factorisation rationale, validation invariants
- Inflow Non-Negativity Solution Methods — Handling of negative inflow realizations from PAR sampling
- SDDP Algorithm — Forward/backward pass structure, cut generation, convergence; references sampling scheme abstraction (sections 3.1–3.2)
- Policy Graphs — The node/transition/pool structure the opening tree’s per-node opening set (section 2.3) is scoped to
- Cut Management — Cut generation, storage, and selection; uses opening tree branchings as the backward pass input
- Multi-Resolution Studies — Multi-resolution modeling configurations; uses scenario generation with varied stage counts
- Post-Study Boundary & Chained Studies — a study chained to an upstream policy; each study draws its own scenario tree
- Determinism & Provenance — bit-identical results per binary at any rank or thread count and on every re-run; grounded in the seed derivation architecture of section 2.2