Training loop¶
Maturity: verified
The training loop owns the ordered transition from a batch to a completed optimizer step.
Complete pretraining loop¶
The current pretraining application owns an explicit train/validation loop
bounded by training.max_steps. It advances a deterministic stateful training
loader, runs full validation at configured step boundaries, applies a fixed
warmup-stable-decay learning-rate schedule, and publishes checkpoint-bound
metric evidence.
One completed step¶
For each global step, the application:
reads the next deterministic stateful training batch;
moves tensor fields to the configured CPU or NVIDIA CUDA GPU, selected by the exact
training.deviceliteralcpuorcuda;normalizes and patches the batch with
prepare_iq_input;runs the next-patch-prediction optimizer step with configured float32 or CUDA bfloat16 arithmetic and optional gradient clipping;
records exact train evidence and runs full validation when due;
saves full mutable state at checkpoint boundaries and at the final step.
The training sampler is deterministically shuffled from the configured seed; validation uses the same complete dataset. The terminal manifest binds the metric prefix and exact latest and best checkpoint references.
The reusable optimizer step requires a finite loss and finite global gradient norm. Checkpoint creation separately rejects non-finite model or optimizer tensors. A successful manifest is published only after the resolved config, frozen metric stream, and terminal checkpoint exist and have byte counts and SHA-256 hashes.
Retry and provider responsibilities¶
run_pretraining_attempt creates exactly one attempt namespace. For resume, it
verifies the parent attempt before creating the child namespace and restores the
model, optimizer, scheduler, loader, loop, metric cursor, and CPU/GPU random
state. Its caller owns retries and cross-attempt orchestration. The application
does not build a container, submit a provider job, monitor that job, or delete
remote resources.