Policy Checkpoint
The policy checkpoint is the directory under policy.path (policy/ by default)
that novomodelo run writes after training. It holds the pool-keyed cuts/NNN.bin
files, the node-keyed basis/NNN.bin files, the stage-keyed states/NNN.bin
trial-point exports (only when exports.states is true), and manifest.bin,
written last.
policy/cuts/NNN.bin
Section titled “policy/cuts/NNN.bin”FlatBuffers binary file encoding all cuts for a single pool. Pool-keyed:
one file per pool (cuts/<pool_id>.bin), zero-padded to three digits (e.g.
000.bin, 012.bin); a pool shared by several leaf nodes appears once. On a
plain stage chain, pool id equals stage index. The general node -> pool mapping
(for a branching policy graph) is resolved through the nodes[].pool_id
entries of policy/manifest.bin, never trusted from the
file name.
Methodology: Cut Management
Each cuts/<pool>.bin is self-describing: alongside its cuts it carries its
own state_dimension, its own cost_scale_factor (the authoritative load-time
scale for this pool — see Cost-Scale Canonicalization),
and its own graph identity (node and graph-stage ids), so no field of the
study-global manifest is needed to interpret a pool’s coefficients.
StageCuts header fields (the root table of each cuts/<pool>.bin):
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
stage_id | uint32 | No | — | Pool id (0-based), the key that names cuts/<pool>.bin. |
state_dimension | uint32 | No | — | State-vector length: the length of every cut’s coefficients and of entity_manifest. |
capacity | uint32 | No | — | Total preallocated cut slots in the pool. |
warm_start_count | uint32 | No | — | Number of leading slots loaded from a previous policy at run start; 0 for a fresh run, except in the terminal pool of a run that loads policy.boundary, where it counts the boundary cuts. |
cuts | [AffinePiece] | No | — | The pool’s cuts, one AffinePiece per cut (fields below). |
active_cut_indices | [uint32] | No | — | Positions in cuts of the cuts currently active in the LP. |
populated_count | uint32 | No | — | Number of filled slots; equals the length of cuts. |
entity_manifest | [EntitySlot] | No | — | Per-slot entity identity, one EntitySlot per state dimension (fields below). |
cost_scale_factor | float64 | No | — | Objective cost-scale factor of the writing study; the pool’s authoritative load-time scale (see Cost-Scale Canonicalization). |
node_id | int32 | No | — | Policy-graph id of the pool’s sole owning node; -1 for a shared pool. |
graph_stage_id | int32 | No | — | Graph-stage id of the stage that owns the pool; -1 when unresolved. No load reads it: a boundary load selects its source pool by priced_state_date. |
priced_state_date | int32 | No | — | The owning stage’s end_date encoded YYYYMMDD (year*10000 + month*100 + day); -2147483648 (the int32 minimum) when not recorded. |
The binary is not human-readable. The logical record structure of each cut
(one AffinePiece per entry in cuts) is:
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
piece_id | uint64 | No | — | Unique identifier for this cut across all iterations. Assigned monotonically by the training loop. The Python load_policy dict calls it cut_id (see Policy Management). |
slot_index | uint32 | No | — | LP row position. Required for checkpoint reproducibility and basis warm-starting. |
iteration | uint32 | No | — | Training iteration that generated this cut. |
forward_pass_index | uint32 | No | — | Forward pass index within the generating iteration. |
intercept | float64 | No | — | Pre-computed cut intercept , where is the state at the generating forward pass node and the aggregated stage value there. |
coefficients | [float64] | No | — | The cut’s slope: one coefficient per state dimension, the sensitivity of the future cost to each incoming state value (see Cut Management). Length equals the state dimension — the number of entries in the embedded entity manifest below, also the pool’s own state_dimension (per-pool, not a policy/manifest.bin field). |
is_active | bool | No | — | Whether this cut is currently active in the LP. Inactive cuts are retained for potential reactivation by the cut selection strategy. |
Each NNN.bin also embeds a per-slot entity manifest — one entry per
state-vector dimension, in canonical cut-coefficient order — recording which
entity each coefficient position belongs to. It is what policy-load
validation checks (see Policy Management):
a policy whose dimensions match the current study by count but bind to different
entities is rejected.
EntitySlot fields (one entity_manifest entry per state dimension):
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
entity_type | EntityType | No | — | byte enum: state-dimension kind — 0 HydroStorage, 1 HydroInflowLag, 2 AnticipatedThermalState, 3 HydroTransitBucket. |
entity_id | int32 | No | — | Owning entity id — for a transit bucket, the downstream hydro. |
subindex | uint32 | No | — | Secondary index within the entity: the inflow-lag order, anticipated-commitment ring slot (counted from stage 0: for the Hold Ring slot ), or transit maturity lag. |
was_active | bool | No | — | Whether the owning entity was operationally active at this stage (excluded from the load-time identity check). |
reference_date | int32 | No | — | The HydroInflowLag reference past-stage start_date, encoded YYYYMMDD (year*10000 + month*100 + day); sentinel -2147483648 (the int32 minimum) for every other family. |
interval_start | int32 | No | — | Inclusive start of the forward-family half-open window, encoded YYYYMMDD: a HydroTransitBucket arrival start or an AnticipatedThermalState delivery start. Sentinel -2147483648 for HydroStorage and HydroInflowLag, and for an AnticipatedThermalState ring slot that holds no live commitment at the pool’s stage or whose delivery lands past the study calendar extended by the post-study stages. |
interval_end | int32 | No | — | Exclusive end of the window, paired with interval_start; reads the sentinel -2147483648 wherever interval_start does. |
Wire readability. A reader parses policy/manifest.bin first, requires the
CBVF file identifier and a format_version of 3, and only then decodes the
payloads. There is no conversion between format versions; see
Versioning policy.
Load admission. Every policy load requires the software and
software_version recorded in manifest.bin to equal the running software’s
name and version exactly (see the version gate).
policy/basis/NNN.bin
Section titled “policy/basis/NNN.bin”FlatBuffers binary file encoding the LP simplex basis checkpoint for a single
stage. Node-keyed: files are named by the policy-graph node’s 0-based
position (basis/NNN.bin; the position is the stage_id field below),
zero-padded to three digits, and exist only for nodes training captured a basis
for; on a stage chain the position is the 0-based stage position, one file per
stage. Bases are always written with the end-of-training
checkpoint. Warm-start, resume, and simulation-only loads read them to
warm-start the stage LP solves, using each basis that fits its stage LP and
leaving the others out (see the
stored-basis gate); boundary
loads read none.
Methodology: LP Warm-Start
The logical record structure (the StageBasis root table) is:
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
stage_id | uint32 | No | — | 0-based position of the policy-graph node (equal to the 0-based stage position on a stage chain), not the declared stage id. |
iteration | uint32 | No | — | Training iteration that produced this basis. |
num_columns | uint32 | No | — | Number of LP columns; equals the column_status length. |
num_rows | uint32 | No | — | Number of LP rows; equals the row_status length. |
column_status | [uint8] | No | — | One canonical basis-status code per LP column (variable). See the encoding note below. |
row_status | [uint8] | No | — | One canonical basis-status code per LP row (constraint). See the encoding note below. |
num_cut_rows | uint32 | No | — | Number of trailing rows in row_status that correspond to cut rows (as opposed to structural constraints). |
The column_status/row_status bytes are the canonical solver-neutral
BasisStatus superset 0..=6 — 0 Lower, 1 Basic, 2 Upper, 3 Zero,
4 Nonbasic, 5 Superbasic, 6 Fixed — not a backend-specific encoding; novomodelo
maps each LP backend (HiGHS default, CLP opt-in) onto this canonical set on
export. num_columns and num_rows are wire-only: they let a generated reader
size the status vectors and are not surfaced by the Python load_policy dict
(which derives the lengths from the vectors themselves).
policy/states/NNN.bin
Section titled “policy/states/NNN.bin”FlatBuffers binary file encoding the visited forward-pass trial points for a
single stage. Stage-keyed: one file per stage (states/<stage_id>.bin),
zero-padded to three digits. Present only when exports.states is true
(default is false). The states/ directory is omitted entirely when disabled.
The file is named by stage position, so on a policy graph with several nodes at
one stage it holds the states of the node with the highest declared id among
them. Like the cut files, each NNN.bin embeds the per-slot entity manifest
describing its state dimensions.
Methodology: Cut Management
Trial points are the state vectors observed at each forward-pass scenario during training. The file is an export of them for diagnostics and analysis; no policy load uses its contents. Cut selection scores cuts against an in-memory record of the visited states that training keeps, not against this file.
StageStates fields (the root table of each states/NNN.bin):
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
stage_id | uint32 | No | — | 0-based stage position in the study (not the declared stage id). |
node_id | int32 | No | — | Policy-graph node identity — the declared node id on a branching graph. Sentinel -1 when absent (a writer that never resolved one, or a buffer that omits the field). Distinct from stage_id the moment a graph carries more than one node per stage. |
state_dimension | uint32 | No | — | Length of each state vector; equals the writing pool’s own state_dimension (per-pool, not a policy/manifest.bin field). |
count | uint32 | No | — | Number of state vectors stored for this stage. |
data | [float64] | No | — | Flat array of count * state_dimension elements, row-major (one state per row). |
entity_manifest | [EntitySlot] | No | — | Per-slot entity identity of each state dimension; state_dimension entries, with the fields of the EntitySlot table under policy/cuts/NNN.bin. |
policy/manifest.bin
Section titled “policy/manifest.bin”The study-global checkpoint manifest: a FlatBuffers root (the
CheckpointManifest table, file_identifier "CBVF") describing the checkpoint
at a high level, written last as the commit signal and read first behind
the format_version gate. Every study-global fact a load needs — the study graph,
the stage count, and the producer provenance — lives here. On the wire,
CheckpointManifest is one flat table: the graph and producer fields sit at its
root, beside a nested season_manifest. The Python bindings group them under
graph_manifest, producer and season_manifest (see
load_policy); the tables below
follow that grouping, with the same field names.
Methodology: Policy Graphs
Python view of the manifest
Section titled “Python view of the manifest”Python dict top-level fields:
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
format_version | integer | No | — | On-disk format version; must equal 3 on read, before any payload is parsed (wire readability). Admission is the separate software and version check. |
software | string | Yes | — | Name of the software that wrote this checkpoint: "novomodelo" from every novomodelo writer, None for a buffer without the field (wire id 20). The version gate compares it, and novomodelo.write_policy_checkpoint stamps it. |
software_version | string | No | — | Version of that software, compared verbatim (wire id 1); every policy load requires it to equal the running version. |
created_at | string | No | — | ISO 8601 timestamp when the checkpoint was written. |
num_stages | integer | No | — | Number of stages the graph manifest spans. Must match the case configuration on resume. |
graph_manifest | object | No | — | Graph manifest: node list, edge list and pool-set size; each node names its pool (nodes[].pool_id). |
producer | object | No | — | Producer-namespaced metadata — the training algorithm’s own recorded state. See below. |
season_manifest | object | No | — | Study-global season / periodic-AR compatibility descriptor: cycle_code (0 monthly, 1 weekly, 2 custom, 255 absent: no season map declared), n_seasons, and per-hydro AR orders (ascending by hydro id). Compared on boundary loads only — see Compatibility requirements; full-FCF loads do not compare it. |
graph_manifest fields:
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
graph_manifest.n_pools | integer | No | — | Number of distinct pools (the pool-set size). |
graph_manifest.nodes[] | array | No | — | Every node, in canonical order, each with its stage and pool. See nodes[] fields below. |
graph_manifest.edges[] | array | No | — | Every directed edge with its transition probability. See edges[] fields below. |
nodes[] fields (one entry per policy-graph node):
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
graph_manifest.nodes[].id | integer | No | — | Node id: the declared node id, or, on a stage chain, the 0-based stage position (not the declared stage id). |
graph_manifest.nodes[].stage_id | integer | No | — | Declared stage id of the stage this node sits at. |
graph_manifest.nodes[].pool_id | integer | No | — | Pool whose payload holds this node’s affine pieces (the node -> pool map: leaf nodes sharing a pool all name the same pool_id). |
edges[] fields (one entry per directed policy-graph edge):
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
graph_manifest.edges[].source_id | integer | No | — | Source node id. |
graph_manifest.edges[].target_id | integer | No | — | Target node id. |
graph_manifest.edges[].probability | number | No | — | Transition probability P(source -> target). |
producer fields:
| Name | Type | Nullable | Units | Description |
|---|---|---|---|---|
producer.completed_iterations | integer | No | — | Number of training iterations completed at checkpoint time. |
producer.final_lower_bound | number | No | — | Lower bound value after the final completed iteration. |
producer.lower_bound_history[] | array | No | — | The lower bound after each recorded iteration, oldest first, the last at completed_iterations; empty when the writer recorded none. Resume restores it. |
producer.best_upper_bound | number | Yes | — | The final completed iteration’s upper bound, if available — the last value, not a min-tracked/observed best. novomodelo run always records it; null only in a checkpoint whose writer recorded none. |
producer.max_iterations | integer | No | — | Maximum iterations configured for the run. |
producer.forward_passes | integer | No | — | Number of forward passes per iteration configured for the run. |
producer.warm_start_cuts | integer | No | — | Largest per-pool warm-start cut count: the maximum of warm_start_counts[], not their sum. 0 for a fresh run without policy.boundary. |
producer.warm_start_counts[] | array | No | — | Per-pool warm-start cut counts, in pool-id order. Supersedes warm_start_cuts for per-pool accuracy when non-empty. |
producer.rng_seed | integer | No | — | RNG seed used by the scenario sampler. Required for reproducibility. |
producer.total_visited_states | integer | No | — | Total number of visited state vectors across all nodes, recorded whenever training keeps the visited-state record — when exports.states is true or a training.cut_selection method is set; 0 when neither is. |
producer.training_block_mode | string | No | — | Block mode the artifact was trained under: the shared lowercase mode ("parallel"/"chronological") when every stage agrees, else "mixed". |
producer.training_block_mode_per_stage[] | array | No | — | Per-study-stage training block modes, in study-stage order. Empty when uniform; populated only for mixed-mode studies. |
producer.cost_scale_factor | number | Yes | — | Objective cost-scale factor the writing study resolved (modeling.cost_scale_factor), recorded here as study-global provenance. The authoritative load-time value is carried per pool in each cuts/<pool>.bin (see the cuts payload above), which marks that pool’s cut coefficients/intercepts as canonical currency units; every load path requires it, and a resolved pool missing it is rejected (“predates self-describing cuts …”). See Cost-Scale Canonicalization for the details. |