Execution

Maturity: verified

Execution separates the application that performs training from the executor that provisions a worker, starts it, and manages the provider lifecycle.

Single-process execution

The current application runs one Python process on either a CPU or an NVIDIA CUDA GPU. Managed launch configurations select one GPU machine, and provider adapters preserve that selection. Provider-managed execution is not the same thing as distributed training.

Application responsibility

The application resolves configuration and mounted gs:// paths, reads the canonical dataset, constructs the selected program, performs bounded optimizer steps, and publishes one attempt’s evidence.

Executor responsibility

The caller or provider adapter owns the mounted Cloud Storage namespace, container lifecycle, source and image provenance inputs, job submission, monitoring, and retry policy. An operator requests cancellation; the provider stops the job and removes its temporary compute and disks.

The committed submission scripts validate the launch YAML, full 40-digit source commit, immutable image digest, safe run and attempt IDs, and paired resume inputs. Each uploads the exact launch YAML with create-only semantics and submits one provider job. Neither waits for success.

Distributed training limitations

There is no DistributedDataParallel, FullyShardedDataParallel, multi-process launcher, rank-aware sampler, collective metric reduction, sharded checkpoint, or elastic restart contract. Documentation and launch configuration must not imply otherwise.

See Also