Skip to content

Training Output

Each section below gives the schema of a file or directory that novomodelo run writes under training/, with the four output dictionary files last; the training/metadata.json, training/model_provenance.json and training/hydro_models.json schemas are on the Metadata Files and Hydro Model Artifacts pages.

Per-iteration convergence log. One row per training iteration. 15 columns.

Methodology: Stopping Rules · Upper Bound Evaluation

NameTypeNullableUnitsDescription
iterationInt32No—Training iteration number (1-based).
lower_boundFloat64No—Lower bound on the risk-adjusted cost after this iteration: the first stage’s risk-adjusted value over its openings with the current cuts (lower bound).
upper_boundFloat64No—Upper bound for this iteration. Under a statistical (sampled) forward pass, the mean over forward-pass scenarios. Under an exact (enumerated) forward pass, the path-weighted expected cost of the enumerated paths; when the risk measure is the same CVaR at every stage, the root value of the nested risk-adjusted recursion instead (see Cut Management — when bounds and certificates hold).
upper_bound_stdFloat64Yes—Sample standard deviation of this iteration’s forward-pass scenario costs. null on every row when upper_bound_kind is "exact".
upper_bound_kindUtf8No—Run-level upper-bound regime, the same on every row: "exact" under enumerated selection, "statistical" under sampled selection.
gap_percentFloat64Yes%Relative gap between lower and upper bounds as a percentage. null when the lower bound is zero or negative.
cuts_addedInt32No—Number of new cuts added to the pool during this iteration’s backward pass.
cuts_removedInt32No—Number of cuts deactivated by the cut selection strategy in this iteration.
cuts_activeInt64No—Total number of active cuts across all stages at the end of this iteration.
time_forward_msInt64NomsWall-clock time spent in the forward pass, in milliseconds.
time_backward_msInt64NomsWall-clock time spent in the backward pass, in milliseconds.
time_total_msInt64NomsTotal wall-clock time for this iteration, in milliseconds.
forward_passesInt32No—Number of forward-pass scenario trajectories evaluated in this iteration.
lp_solvesInt64No—Total number of LP solves across all stages and forward passes in this iteration.
mean_rows_in_lpFloat64No—Mean number of resident cut rows loaded per LP solve under dynamic cut selection in this iteration; 0 when no dynamic cut selection ran.

Per-iteration wall-clock timing breakdown by phase. 19 columns. Emitted as one row per (iteration, rank) for rank-only sequential values (worker_id is NULL) and one row per (iteration, rank, worker_id) for per-worker parallel-region values; SUM(col) GROUP BY iteration recovers the per-iteration total for each timing column. rank and worker_id are nullable Int32; the 16 timing columns are non-nullable.

The top-level non-overlapping phases are: forward_wall_ms, backward_wall_ms, cut_selection_ms, mpi_allreduce_ms, and lower_bound_ms. The backward parallel overhead is decomposed into three components: bwd_setup_ms (aggregate non-solve work summed across workers), bwd_load_imbalance_ms (max-worker minus average-worker), and bwd_scheduling_overhead_ms (parallel wall minus max-worker). The forward pass carries the same three sub-components with fwd_ prefix. The backward phase also has the sub-components cut_sync_ms, state_exchange_ms, and cut_batch_build_ms. The residual not attributed to any phase is overhead_ms.

NameTypeNullableUnitsDescription
iterationInt32No—Training iteration number (1-based).
rankInt32Yes—MPI rank that produced this row. Nullable in the schema; always set in written output.
worker_idInt32Yes—Worker index within the rank’s thread pool. NULL for rank-only sequential rows.
forward_wall_msInt64NomsWall-clock time for the forward pass (all stages and scenarios).
backward_wall_msInt64NomsWall-clock time for the backward pass (all stages and trial points).
cut_selection_msInt64NomsTime spent running the cut selection pass.
mpi_allreduce_msInt64NomsTime spent in the forward-pass bound synchronization: an allgatherv of the per-trajectory costs across ranks and their canonical-order reduction.
cut_sync_msInt64NomsTime spent in per-stage cut sync allgatherv (sub-component of backward).
lower_bound_msInt64NomsTime spent evaluating the lower bound (stage-0 LP solves for all openings).
state_exchange_msInt64NomsTime spent in state exchange allgatherv (sub-component of backward).
cut_batch_build_msInt64NomsTime spent assembling cut row batches (sub-component of backward).
bwd_setup_msInt64NomsAggregate non-solve work (model load, bound updates and basis installation) summed across backward workers, in ms. May exceed backward_wall_ms; it is a cost metric, not a wall-time slice.
bwd_load_imbalance_msInt64NomsBackward load imbalance: max_worker_total - avg_worker_total, clamped to zero.
bwd_scheduling_overhead_msInt64NomsBackward scheduling overhead: parallel_wall - max_worker_total, clamped to zero.
fwd_setup_msInt64NomsAggregate non-solve work summed across forward workers, in ms. Same aggregate semantics as bwd_setup_ms.
fwd_load_imbalance_msInt64NomsForward load imbalance: max_worker_total - avg_worker_total, clamped to zero.
fwd_scheduling_overhead_msInt64NomsForward scheduling overhead: parallel_wall - max_worker_total, clamped to zero.
overhead_msInt64NomsResidual wall-clock time not attributed to any of the above phases.
lazy_scoring_msInt64NomsPer-worker time spent in lazy candidate scoring inside the lazy-selection solve. A sub-component of the forward/backward phases (not a top-level addend); 0 when the lazy path is unused.

Per-iteration, per-phase, per-stage, per-opening, per-worker LP solver statistics for diagnosing conditioning issues and retry behavior. One row per (iteration, phase, stage_id, opening_index, rank, worker_id) tuple on the backward phase (per-opening, per-worker); one row per (iteration, phase, stage_id) tuple on the forward, lower_bound, and simulation phases. A training row fills iteration and leaves scenario_id NULL; a simulation row fills scenario_id and leaves iteration NULL. stage_id is NULL on lower_bound rows (no stage); opening_index and worker_id are NULL wherever the row has no per-opening/per-worker dimension, and rank is NULL on a simulation row only. 19 columns. iteration, scenario_id, stage_id, opening_index, rank, and worker_id are nullable Int32; all other columns are non-nullable.

Methodology: LP Warm-Start

NameTypeNullableUnitsDescription
iterationInt32Yes—Training iteration (1-based). NULL on a simulation row (scenario_id is filled instead).
scenario_idInt32Yes—Simulation scenario id (0-based). NULL on a training row (iteration is filled instead).
phaseUtf8No—"forward", "backward", "lower_bound", or "simulation".
stage_idInt32Yes—Declared stage id (from stages.json). NULL on lower_bound rows.
opening_indexInt32Yes—Opening (noise realization) index within the stage for backward rows. NULL for forward, lower_bound, simulation.
rankInt32Yes—MPI rank that produced this row. NULL on a simulation row.
worker_idInt32Yes—Worker index within the rank’s thread pool. NULL for rows without a per-worker dimension.
lp_solvesUInt32No—Number of LP solves in this row’s bucket.
lp_successesUInt32No—Number of solves that returned optimal.
lp_retriesUInt32No—Number of solves that required at least one retry.
lp_failuresUInt32No—Number of solves that failed after exhausting all retry levels.
retry_attemptsUInt32No—Total retry attempts across all LP solves in this bucket.
basis_offeredUInt32No—Number of solves offered a starting basis (warm-start attempts).
basis_consistency_failuresUInt32No—Number of warm-start solves whose basis the solver rejected as inconsistent with the model. A basis with fewer rows than the model is rejected before it reaches the solver; such a solve counts here but not in basis_offered.
simplex_iterationsUInt64No—Total simplex iterations (or IPM iterations) across all solves.
solve_time_msFloat64NomsCumulative LP solve wall-clock time in milliseconds.
load_model_time_msFloat64NomsCumulative time spent loading the stage model into the solver, in milliseconds.
set_bounds_time_msFloat64NomsCumulative time spent updating row and column bounds, in milliseconds.
basis_set_time_msFloat64NomsCumulative time spent installing bases for warm-start, in milliseconds.

Per-level retry success counts, normalized from the solver iterations table. One row per (iteration, phase, stage_id, retry_level) tuple where the count is positive (sparse encoding). Rows cover lower_bound solves and, under enumerated training selection, backward solves. forward solves, and backward solves under sampled selection, add no rows; their retries appear only in lp_retries and retry_attempts of training/solver/iterations.parquet. 5 columns. All non-nullable except stage_id.

NameTypeNullableUnitsDescription
iterationUInt32No—Training iteration number (1-based).
phaseUtf8No—Algorithm phase: "backward" or "lower_bound".
stage_idInt32Yes—Declared stage id (from stages.json) on backward rows; null on lower_bound rows.
retry_levelUInt32No—Retry escalation level: 0 to 11 under HiGHS; the CLP backend records no per-level counts, so a CLP run writes no rows. See the Solver Safeguards section of the Performance Accelerators guide.
countUInt64No—Number of LP solves recovered at this level. Counts are not summed across ranks or workers.

LP prescaling diagnostics written once after stage template construction. Documents the coefficient ranges before and after column/row scaling for each stage, plus the applied scale-factor distributions. Useful for diagnosing numerical conditioning issues.

Methodology: LP Layout and Scaling

The JSON is a single top-level object:

{
"cost_scale_factor": 1000000.0,
"stages": [
{
"stage_id": 0,
"dimensions": { "num_cols": 128, "num_rows": 96, "num_nz": 412 },
"pre_scaling": {
"matrix_coeff_range": [0.001, 5000.0],
"matrix_coeff_ratio": 5000000.0,
"objective_range": [1.0, 250000.0],
"objective_ratio": 250000.0
},
"post_scaling": {
"matrix_coeff_range": [0.5, 4.2],
"matrix_coeff_ratio": 8.4,
"objective_range": [0.8, 3.1],
"objective_ratio": 3.875
},
"col_scale": { "min": 0.02, "max": 48.0, "median": 1.0, "count": 128 },
"row_scale": { "min": 0.1, "max": 12.0, "median": 1.0, "count": 96 }
}
],
"summary": {
"worst_pre_scaling_matrix_ratio": 5000000.0,
"worst_post_scaling_matrix_ratio": 8.4,
"improvement_factor": 595238.1,
"num_stages": 1
}
}

Top-level fields:

NameTypeNullableUnitsDescription
cost_scale_factornumberNo—Cost scale factor applied to objective coefficients during template build.
stagesarrayNo—One entry per stage. See “stages[] fields” below.
summaryobjectNo—Cross-stage summary. See “summary fields” below.

stages[] fields:

NameTypeNullableUnitsDescription
stages[].stage_idintegerNo—0-based stage position in the study (not the declared stage id).
stages[].dimensionsobjectNo—LP dimensions for this stage’s template. See “dimensions fields” below.
stages[].pre_scalingobjectNo—Coefficient ranges before column/row scaling. See “pre_scaling / post_scaling fields” below.
stages[].post_scalingobjectNo—Coefficient ranges after column/row (and cost) scaling. Same shape as pre_scaling.
stages[].col_scaleobjectNo—Summary of the column scale-factor vector. See “col_scale / row_scale fields” below.
stages[].row_scaleobjectNo—Summary of the row scale-factor vector. Same shape as col_scale.

dimensions fields:

NameTypeNullableUnitsDescription
stages[].dimensions.num_colsintegerNo—Number of columns (decision variables).
stages[].dimensions.num_rowsintegerNo—Number of structural rows (constraints).
stages[].dimensions.num_nzintegerNo—Number of nonzero entries in the constraint matrix.

pre_scaling / post_scaling fields:

NameTypeNullableUnitsDescription
stages[].pre_scaling.matrix_coeff_rangearrayNo—Array of two numbers: [min, max] absolute value over nonzero constraint-matrix entries.
stages[].pre_scaling.matrix_coeff_rationumberNo—Ratio of the largest to smallest absolute nonzero matrix coefficient (max / min).
stages[].pre_scaling.objective_rangearrayNo—Array of two numbers: [min, max] absolute value over nonzero objective coefficients.
stages[].pre_scaling.objective_rationumberNo—Ratio of the largest to smallest absolute nonzero objective coefficient (max / min).

col_scale / row_scale fields:

NameTypeNullableUnitsDescription
stages[].col_scale.minnumberNo—Minimum scale factor.
stages[].col_scale.maxnumberNo—Maximum scale factor.
stages[].col_scale.mediannumberNo—Median scale factor.
stages[].col_scale.countintegerNo—Number of scale factors (num_cols for col_scale, num_rows for row_scale).

summary fields:

NameTypeNullableUnitsDescription
summary.worst_pre_scaling_matrix_rationumberNo—Maximum pre-scaling matrix coefficient ratio across all stages.
summary.worst_post_scaling_matrix_rationumberNo—Maximum post-scaling matrix coefficient ratio across all stages.
summary.improvement_factornumberNo—worst_pre_scaling_matrix_ratio / worst_post_scaling_matrix_ratio.
summary.num_stagesintegerNo—Number of stages.

Per-stage cut selection statistics. One row per (iteration, stage_id) pair, written only at iterations where selection ran. 10 columns.

Methodology: Cut Management

NameTypeNullableUnitsDescription
iterationInt32No—Training iteration number (1-based).
stage_idInt32No—Declared stage id (from stages.json).
cuts_populatedInt32No—Total cut slots containing cuts (active + inactive).
cuts_active_beforeInt32No—Active cuts before this iteration’s selection pass.
cuts_deactivatedInt32No—Cuts deactivated by the selection pass.
cuts_reactivatedInt32No—Cuts reactivated by the selection pass.
cuts_active_afterInt32No—Active cuts after the selection pass.
selection_time_msFloat64NomsWall-clock time of this stage’s selection pass.
budget_evictedInt32Yes—Cuts evicted by the budget pass. null when max_active_per_stage is not set.
active_after_budgetInt32Yes—Active cuts after the budget pass. null when max_active_per_stage is not set.

Four self-documenting files that allow output Parquet files to be interpreted without reference to the original input case. All files are written atomically.

Static mapping from integer codes to human-readable labels for all categorical fields used in Parquet output. The same mapping applies for the lifetime of a release (the version field tracks breaking changes).

{
"version": "1.0",
"generated_at": "<timestamp>",
"operative_state": {
"0": "deactivated",
"1": "maintenance",
"2": "operating",
"3": "saturated"
},
"storage_binding": {
"0": "none",
"1": "below_minimum",
"2": "above_maximum",
"3": "both"
},
"contract_type": {
"0": "import",
"1": "export"
},
"entity_type": {
"0": "hydro",
"1": "thermal",
"2": "bus",
"3": "line",
"4": "pumping_station",
"5": "contract",
"7": "non_controllable",
"8": "hydro_unit_group"
},
"bound_type": {
"0": "storage_min",
"1": "storage_max",
"2": "turbined_min",
"3": "turbined_max",
"4": "outflow_min",
"5": "outflow_max",
"6": "generation_min",
"7": "generation_max",
"8": "flow_min",
"9": "flow_max"
}
}

One row per entity across all entity types, plus one row per hydro unit group (entity type code 8). Columns:

NameTypeNullableUnitsDescription
entity_type_codeintegerNo—Integer entity type code (see codes.json entity_type mapping).
entity_idintegerNo—Integer entity ID matching the *_id column in the corresponding simulation Parquet file. For a hydro unit group row, this is the group’s id, which is scoped to its plant, not global.
namestringNo—Human-readable entity name from the case input files. A hydro unit group row’s name is "{hydro_id}/{group_name}", plant-qualified since the group id alone is not globally unique.
bus_idintegerNo—Integer bus ID to which this entity is connected. For buses, equals entity_id. -1 for a line (connects two buses) and for a hydro (the plant’s unit groups own the bus association; a split plant has no single owning bus) — a hydro unit group row carries that group’s own bus_id instead.
system_idintegerNo—System partition index. Always 0 (single-system cases).

Rows are ordered by entity_type_code ascending, then by canonical entity order within each type (operational_start_date, then entity_id) — except type code 8: a group’s entity_id is plant-scoped, so those rows order plant-major (canonical hydro order), then group-minor (each plant’s own id-sorted unit_groups order).

One row per column of every output schema that carries a variables.csv label. The file column holds that label, not a file path; the label table below maps each label to its output file. Documents every column’s name, type, unit of measure, description, and nullability. Useful for building generic result readers that do not hard-code column names.

NameTypeNullableUnitsDescription
filestringNo—Label of the output schema this column belongs to (e.g. "hydros", "costs"); not a file path. The label table lists the files for each label.
columnstringNo—Exact column name as it appears in the Parquet file.
typestringNo—Lowercase column type token: i8, i32, i64, u32, u64, f64, bool, string, date32, or unknown.
unitstringNo—Physical unit, spelled in ASCII: "" for dimensionless, code, boolean and id columns, or one of $, $/MWh, $/hm3, %, MW, MW/(m3/s), MWh, hm3, m3/s, ms; varies when the unit depends on the row (generic_violations.slack_value).
descriptionstringNo—Short description of the column’s meaning.
nullablestringNo—"true" or "false".

The table below maps each file label to its output file or directory.

LabelOutput file
costssimulation/costs/
hydrossimulation/hydros/
hydro_bus_generationsimulation/hydro_bus_generation/
thermalssimulation/thermals/
exchangessimulation/exchanges/
busessimulation/buses/
pumping_stationssimulation/pumping_stations/
contractssimulation/contracts/
non_controllablessimulation/non_controllables/
inflow_lagssimulation/inflow_lags/
in_transitsimulation/in_transit/
transit_seedsimulation/transit_seed/
generic_violationssimulation/violations/generic/
pathssimulation/paths.parquet
scenario_summarysimulation/scenario_summary.parquet
convergencetraining/convergence.parquet
iteration_timingtraining/timing/iterations.parquet
cut_selectiontraining/cut_selection/iterations.parquet
solver_iterationstraining/solver/iterations.parquet and simulation/solver/iterations.parquet
retry_histogramtraining/solver/retry_histogram.parquet and simulation/solver/retry_histogram.parquet

variables.csv has no rows for simulation/anticipated_lanes/, anticipated/fixed_deliveries.parquet, generic_constraints/resolved_echo.parquet, training/dictionaries/bounds.parquet, or the Parquet files under hydro_models/ and stochastic/.

Per-entity, per-stage resolved bounds: each row holds the entity’s value after stage overrides, not a penalty. Each (entity, stage, bound type) has one stage-level row with block_id null, plus one row with block_id set for each block that a per-block override resolves. Buses and non-controllable sources have no rows. The list below gives each entity family, named by its codes.json entity_type label, with its bound types:

  • hydro: storage_min, storage_max, turbined_min, turbined_max, outflow_min, outflow_max, generation_min, generation_max. The stage-level outflow_max row exists only for a plant that has a maximum outflow. Block rows cover the turbined, outflow, and generation bounds.
  • thermal: generation_min, generation_max, with stage-level and block rows.
  • line: flow_min, always 0, and flow_max, the direct capacity. Block rows cover the direct capacity only; the reverse capacity is not reported.
  • pumping_station and contract: flow_min, flow_max, with stage-level and block rows.
  • hydro_unit_group: turbined_min, turbined_max, generation_min, generation_max, with stage-level and block rows. hydro_id holds the owning plant’s id, because a group’s entity_id is plant-scoped.

Only the ten bound types in codes.json are written. The line reverse capacity, the contract price, and the hydro diversion maximum have no rows, at the stage level or per block.

NameTypeNullableUnitsDescription
entity_type_codeInt8No—Entity type code (see codes.json).
entity_idInt32No—Entity ID.
hydro_idInt32Yes—Owning plant id on a hydro unit group row (entity_type_code 8), whose entity_id is plant-scoped; null on every other row.
stage_idInt32No—Declared stage id from stages.json (not a position).
block_idInt32Yes—0-based block index on a per-block override row; null on the stage-level row.
bound_type_codeInt8No—Bound type code (see codes.json bound_type mapping).
bound_valueFloat64No—Resolved bound value in the bound’s natural unit.