Cloud Batch post-training¶
Maturity: runtime-isolated
The Cloud Batch adapter starts the same radio-frequency (RF) classifier post-training command used locally. It owns the provider launch configuration, resource selection, submission, and provider job identity. The application still owns training, validation-based checkpoint selection, immutable artifact creation, procedure-controlled test evaluation, and terminal manifest publication. Validation selection does not construct or iterate the test dataset; final qualification evaluates the selected artifact once on the untouched test split.
Launch inputs¶
Run infra/runtime/submit_cloud_batch_posttraining.sh from the repository root.
The adapter requires:
JOB_ID: a lowercase Cloud Batch job IDIMAGE_URI: an immutable runtime image ending in@sha256:<64 hex digits>POSTTRAIN_CONFIG_PATH: the resolved application YAML path inside the mounted runs bucketLAUNCH_CONFIG_PATH: a readable Cloud Batch launch configuration, normallyconfigs/launch/cloud_batch_l4_dataset_cache.yamlLAUNCH_CONFIG_URI: the create-only destination for that exact launch file below the attempt’slaunch/directory in the runs bucket
For example:
export JOB_ID=<cloud-batch-job-id>
export IMAGE_URI=<region>-docker.pkg.dev/<project>/<repository>/<image>@sha256:<digest>
export POSTTRAIN_CONFIG_PATH=/mnt/disks/runs/<runs-bucket>/configs/posttrain.yaml
export LAUNCH_CONFIG_PATH=configs/launch/cloud_batch_l4_dataset_cache.yaml
export LAUNCH_CONFIG_URI=gs://<runs-bucket>/<run-id>/attempts/<attempt-id>/launch/cloud_batch_l4_dataset_cache.yaml
infra/runtime/submit_cloud_batch_posttraining.sh
The post-training YAML supplies the classification run and attempt IDs, the
required validation_selection or final_qualification procedure, dataset and
checkpoint inputs, training settings, application output location, and source
revision. Its dataset and output paths must use the dataset and runs
mounts selected by the launch configuration. Leave container_image_uri,
provider_job_id, launch_config_uri, and launch_config_sha256 unset in a
managed recipe so the provider adapter supplies the identities of this launch.
Submission behavior¶
Before creating the job, the adapter validates the single-worker launch shape and uploads the exact launch YAML with create-only Cloud Storage semantics. It computes that file’s SHA-256 and passes the following provider identities to the application:
immutable container image URI
Cloud Batch job ID
launch-config URI
launch-config SHA-256
The application records those identities in its terminal run_manifest.json.
The adapter never edits that manifest.
The job runs exactly:
python -m rffm.cli.main posttrain \
--config POSTTRAIN_CONFIG_PATH \
--hidden-state-cache-directory /mnt/disks/rffm-cache/posttraining-hidden-states
The launch configuration selects one GPU machine type and boot image, and the adapter passes both values to Cloud Batch unchanged. The checked-in launch file is one concrete operational configuration, not a machine or image allowlist. Every alternate configuration requires its own operational qualification.
The boot image must satisfy the package-installation choices in that file. With
install_gpu_drivers: true, the image must support Batch-managed NVIDIA driver
installation; with false, it must already provide a driver compatible with
the selected GPU. With install_ops_agent: true, the image must support
Batch-managed Ops Agent installation; with false, it must provide the agent
when those metrics are required. Post-training requires
gpu_profiling: disabled; the optional DCGM dependencies documented for
pretraining do not apply.
Application-level PyTorch profiling is separate from the launch file’s
gpu_profiling setting. An operator can add the following optional block to a
post-training recipe with at least two epochs:
profiling:
batch_count: 10
The post-training command must still receive
--hidden-state-cache-directory. It then writes standard Chrome traces and a
cache_lifecycle_summary.json under the attempt’s application/profiling/
directory. The bounded epoch-one traces show source-batch wait, input
preparation, frozen-backbone execution, cache-store calls, classifier work, and
the split-boundary wait for background writes. Epoch-two traces show the first
completed-cache reuse, the one-batch-ahead host-to-GPU prefetch, and classifier
work. Training and validation are traced separately.
The summary adds durations measured inside DataLoader workers for raw and mmap
reads and inside the cache caller/background writer for preallocation, staging,
backpressure, CUDA-copy waits, and actual mmap assignments. The read sample is
bounded by batch_count; write measurements span the complete epoch-one cache
construction. These inclusive timings expose the I/O lifecycle but do not
attribute individual filesystem system calls or page faults. The files are
attempt-scoped diagnostics rather than terminal-manifest evidence and may exist
for an interrupted attempt. Profiler overhead means their durations should
guide bottleneck analysis, not serve as benchmark measurements.
Phase durations are inclusive and non-additive because worker observations can run concurrently and writer-task totals contain nested phases. Use backpressure and completion waits as critical-path signals; do not sum phase totals into an epoch duration or expected runtime reduction.
The job disables retries and mounts a temporary local solid-state drive for scratch storage, with its capacity selected by the launch configuration. Cloud Storage FUSE uses that drive for its read-only dataset cache. During the first complete training and validation passes, the application also writes BF16 frozen-backbone hidden states there. It restores FP32 before classifier work; later epochs read the same mapped BF16 tensors without reopening signal records or running the backbone. The runs bucket remains mounted read-write. Hidden states are never copied to the runs bucket, listed in the terminal manifest, or retained after Cloud Batch releases the worker and temporary disk.
The adapter and generated job definition are validated locally. Actual managed training, artifact publication, and cleanup evidence require separate operational approval.