Execution API

Maturity: verified

Execution APIs define how callers start a unit of work and inspect its result without making provider lifecycle part of the application contract.

Pretraining interfaces

The current surface executes one complete fresh or resumed pretraining attempt. The earlier single-loader executor and partial model/optimizer checkpoint format have been removed; the application now has one training and evidence path.

Application lifecycle

Symbol

Contract

PretrainAttemptContext

caller-supplied run, attempt, repository, mount, and optional provider provenance

run_pretraining_attempt

execute the application-owned train/validation loop and publish terminal evidence

PretrainAttemptRecord

successful attempt identities, artifact URIs, final step, and loss

Checkpoints and evidence

Symbol

Contract

TrainValidationEngine

advance one completed training step and any due full-validation boundary

NPPMetricEvent

exact train/validation metric-event vocabulary

MetricSegmentWriter

append metric events and bind checkpoint-safe byte prefixes

save_full_training_checkpoint

save model, optimizer, scheduler, loader, loop, random, and metric-prefix state

PretrainCheckpointReference

reuse one checkpoint already verified by the checkpoint writer in terminal evidence

write_success_manifest

bind the frozen metric cursor and exact latest/best checkpoints into one immutable result

verify_resume_source

verify the metric records and checkpoint selections used for exact resume

The executor performs full validation at configured boundaries, follows the fixed warmup-stable-decay learning-rate schedule, and records exact latest and best checkpoint references. On resume, it verifies the prior attempt before creating the child workspace, restores the full PyTorch and loader state, copies the checkpoint-bound metric prefix, and continues with a new attempt identity. The caller still owns retry and provider lifecycle; distributed execution is outside this contract.

See Also