Execution¶
Maturity: verified
Execution separates the application that performs training from the executor that provisions a worker, starts it, and manages the provider lifecycle.
Single-process execution¶
The current application runs one Python process on either a CPU or an NVIDIA CUDA GPU. Managed launch configurations select one GPU machine, and provider adapters preserve that selection. Provider-managed execution is not the same thing as distributed training.
Application responsibility¶
The application resolves configuration and mounted gs:// paths, reads the
canonical dataset, constructs the selected program, performs bounded optimizer
steps, and publishes one attempt’s evidence.
Executor responsibility¶
The caller or provider adapter owns the mounted Cloud Storage namespace, container lifecycle, source and image provenance inputs, job submission, monitoring, and retry policy. An operator requests cancellation; the provider stops the job and removes its temporary compute and disks.
The committed submission scripts validate the launch YAML, full 40-digit source commit, immutable image digest, safe run and attempt IDs, and paired resume inputs. Each uploads the exact launch YAML with create-only semantics and submits one provider job. Neither waits for success.
Distributed training limitations¶
There is no DistributedDataParallel, FullyShardedDataParallel, multi-process launcher, rank-aware sampler, collective metric reduction, sharded checkpoint, or elastic restart contract. Documentation and launch configuration must not imply otherwise.