Cloud Batch post-training

Maturity: runtime-isolated

The Cloud Batch adapter starts the same radio-frequency (RF) classifier post-training command used locally. It owns the provider launch configuration, resource selection, submission, and provider job identity. The application still owns training, validation-based checkpoint selection, immutable artifact creation, procedure-controlled test evaluation, and terminal manifest publication. Validation selection does not construct or iterate the test dataset; final qualification evaluates the selected artifact once on the untouched test split.

Launch inputs

Run infra/runtime/submit_cloud_batch_posttraining.sh from the repository root. The adapter requires:

  • JOB_ID: a lowercase Cloud Batch job ID

  • IMAGE_URI: an immutable runtime image ending in @sha256:<64 hex digits>

  • POSTTRAIN_CONFIG_PATH: the resolved application YAML path inside the mounted runs bucket

  • LAUNCH_CONFIG_PATH: a readable Cloud Batch launch configuration, normally configs/launch/cloud_batch_l4_dataset_cache.yaml

  • LAUNCH_CONFIG_URI: the create-only destination for that exact launch file below the attempt’s launch/ directory in the runs bucket

For example:

export JOB_ID=<cloud-batch-job-id>
export IMAGE_URI=<region>-docker.pkg.dev/<project>/<repository>/<image>@sha256:<digest>
export POSTTRAIN_CONFIG_PATH=/mnt/disks/runs/<runs-bucket>/configs/posttrain.yaml
export LAUNCH_CONFIG_PATH=configs/launch/cloud_batch_l4_dataset_cache.yaml
export LAUNCH_CONFIG_URI=gs://<runs-bucket>/<run-id>/attempts/<attempt-id>/launch/cloud_batch_l4_dataset_cache.yaml

infra/runtime/submit_cloud_batch_posttraining.sh

The post-training YAML supplies the classification run and attempt IDs, the required validation_selection or final_qualification procedure, dataset and checkpoint inputs, training settings, application output location, and source revision. Its dataset and output paths must use the dataset and runs mounts selected by the launch configuration. Leave container_image_uri, provider_job_id, launch_config_uri, and launch_config_sha256 unset in a managed recipe so the provider adapter supplies the identities of this launch.

Submission behavior

Before creating the job, the adapter validates the single-worker launch shape and uploads the exact launch YAML with create-only Cloud Storage semantics. It computes that file’s SHA-256 and passes the following provider identities to the application:

  • immutable container image URI

  • Cloud Batch job ID

  • launch-config URI

  • launch-config SHA-256

The application records those identities in its terminal run_manifest.json. The adapter never edits that manifest.

The job runs exactly:

python -m rffm.cli.main posttrain \
  --config POSTTRAIN_CONFIG_PATH \
  --hidden-state-cache-directory /mnt/disks/rffm-cache/posttraining-hidden-states

The launch configuration selects one GPU machine type and boot image, and the adapter passes both values to Cloud Batch unchanged. The checked-in launch file is one concrete operational configuration, not a machine or image allowlist. Every alternate configuration requires its own operational qualification.

The boot image must satisfy the package-installation choices in that file. With install_gpu_drivers: true, the image must support Batch-managed NVIDIA driver installation; with false, it must already provide a driver compatible with the selected GPU. With install_ops_agent: true, the image must support Batch-managed Ops Agent installation; with false, it must provide the agent when those metrics are required. Post-training requires gpu_profiling: disabled; the optional DCGM dependencies documented for pretraining do not apply.

Application-level PyTorch profiling is separate from the launch file’s gpu_profiling setting. An operator can add the following optional block to a post-training recipe with at least two epochs:

profiling:
  batch_count: 10

The post-training command must still receive --hidden-state-cache-directory. It then writes standard Chrome traces and a cache_lifecycle_summary.json under the attempt’s application/profiling/ directory. The bounded epoch-one traces show source-batch wait, input preparation, frozen-backbone execution, cache-store calls, classifier work, and the split-boundary wait for background writes. Epoch-two traces show the first completed-cache reuse, the one-batch-ahead host-to-GPU prefetch, and classifier work. Training and validation are traced separately.

The summary adds durations measured inside DataLoader workers for raw and mmap reads and inside the cache caller/background writer for preallocation, staging, backpressure, CUDA-copy waits, and actual mmap assignments. The read sample is bounded by batch_count; write measurements span the complete epoch-one cache construction. These inclusive timings expose the I/O lifecycle but do not attribute individual filesystem system calls or page faults. The files are attempt-scoped diagnostics rather than terminal-manifest evidence and may exist for an interrupted attempt. Profiler overhead means their durations should guide bottleneck analysis, not serve as benchmark measurements.

Phase durations are inclusive and non-additive because worker observations can run concurrently and writer-task totals contain nested phases. Use backpressure and completion waits as critical-path signals; do not sum phase totals into an epoch duration or expected runtime reduction.

The job disables retries and mounts a temporary local solid-state drive for scratch storage, with its capacity selected by the launch configuration. Cloud Storage FUSE uses that drive for its read-only dataset cache. During the first complete training and validation passes, the application also writes BF16 frozen-backbone hidden states there. It restores FP32 before classifier work; later epochs read the same mapped BF16 tensors without reopening signal records or running the backbone. The runs bucket remains mounted read-write. Hidden states are never copied to the runs bucket, listed in the terminal manifest, or retained after Cloud Batch releases the worker and temporary disk.

The adapter and generated job definition are validated locally. Actual managed training, artifact publication, and cleanup evidence require separate operational approval.

See Also