Cloud Batch GPU training with dataset cache

Maturity: runtime-isolated

Cloud Batch is the single-worker training path when canonical dataset shards need the standard Cloud Storage FUSE file cache on local disk.

Launch configuration

infra/runtime/submit_cloud_batch_training.sh reads configs/launch/cloud_batch_l4_dataset_cache.yaml, uploads that exact file, and submits one Batch job. The launch configuration selects one GPU machine type and boot image, and the adapter passes both values to Cloud Batch unchanged. The selected machine must provide the GPU resources required by the application. The checked-in launch file is one concrete operational configuration, not a machine or image allowlist. Every alternate configuration requires its own operational qualification.

The boot image must satisfy the package-installation choices in the launch file. With install_gpu_drivers: true, the image must support Batch-managed NVIDIA driver installation; with false, it must already provide a driver compatible with the selected GPU. With install_ops_agent: true, the image must support Batch-managed Ops Agent installation; with false, it must provide the agent when those metrics are required. The checked-in configuration uses Google’s provider-managed batch-debian image-family alias and enables both installers. The optional DCGM setup remains tied to its exact Debian 11 AMD64 package and systemd assumptions. The 100 GB balanced persistent boot disk stores the operating system, container, and runtime files. The temporary 375 GB local SSD is used only by the dataset file cache.

gpu_profiling selects disabled or dcgm and defaults to disabled. Standard CPU, GPU, memory, disk, and network metrics require the configured driver and Ops Agent installation paths to succeed. Select dcgm only when GPU occupancy and execution-pipe detail is required. Any image used with dcgm must satisfy the setup script’s exact Debian 11 AMD64 package and systemd assumptions.

The script requires JOB_ID, RUN_ID, ATTEMPT_ID, EXPERIMENT, IMAGE_URI, SOURCE_REVISION, LAUNCH_CONFIG_PATH, and LAUNCH_CONFIG_URI. It accepts the same optional CONFIG_OVERRIDES_PATH, RESUME_FROM, and TARGET_MAX_STEPS inputs as the Vertex adapter.

CONFIG_OVERRIDES_PATH names a text file with one application setting per line. For example, a job can select a different device batch size and periodic checkpoint interval without changing the Batch launch configuration:

input_pipeline.batch_size=32
checkpoint.save_every_n_steps=500

The application validates the resolved configuration and retains it with the run artifacts.

Storage behavior

Batch mounts the canonical dataset bucket read-only. The mount uses Cloud Storage FUSE’s cache-dir and file-cache-cache-file-for-range-read options, with the cache directory on the local SSD. A separate uncached writable mount holds run artifacts. Cloud Storage FUSE owns all downloading, caching, and eviction; the repository provides no cache implementation.

Operational metrics

In the committed configuration, Batch installs the NVIDIA driver and Google Cloud Ops Agent on the temporary Debian worker. Compute Engine and the Ops Agent then publish standard CPU, host-memory, local-disk, network, and NVIDIA GPU measurements in every mode. Alternate images must supply equivalent capabilities through their configured installation source.

When gpu_profiling: dcgm is selected, a host setup runnable downloads the SHA-256-verified NVIDIA Data Center GPU Manager (DCGM) 3.3.9 package, starts the privileged DCGM service with its unused system-monitoring module disabled, configures the Ops Agent’s version-1 DCGM receiver, and verifies all three components. Training does not start unless the same DCGM host-engine process remains active and responsive. Useful metric types include:

  • compute.googleapis.com/instance/cpu/utilization

  • compute.googleapis.com/instance/network/received_bytes_count and compute.googleapis.com/instance/network/sent_bytes_count

  • agent.googleapis.com/memory/percent_used

  • agent.googleapis.com/disk/percent_used, agent.googleapis.com/disk/read_bytes_count, and agent.googleapis.com/disk/write_bytes_count

  • agent.googleapis.com/gpu/utilization and agent.googleapis.com/gpu/memory/bytes_used

  • agent.googleapis.com/gpu/processes/utilization and agent.googleapis.com/gpu/processes/max_bytes_used

  • workload.googleapis.com/dcgm.gpu.profiling.sm_occupancy, sm_utilization, dram_utilization, and pipe_utilization

  • workload.googleapis.com/dcgm.gpu.profiling.pcie_traffic_rate and nvlink_traffic_rate when the hardware provides them

The built-in agent.googleapis.com measurements are free. Optional DCGM measurements use chargeable workload.googleapis.com metric types. At the 60-second cadence and one GPU per temporary worker, their ingestion is small compared with GPU compute cost. DCGM uses GPU hardware counters with low overhead; do not run its profiling receiver concurrently with an interactive NVIDIA profiler such as Nsight. The pinned DCGM package is about 911 MB. It adds provisioning time and network transfer only to jobs that select gpu_profiling: dcgm.

The dataset mount also enables Cloud Storage FUSE’s Cloud Monitoring exporter at a 60-second interval. Its file_cache/read_count metric distinguishes cache hits and misses. file_cache/read_bytes_count, file_cache/read_latencies, gcs/download_bytes_count, gcs/read_bytes_count, gcs/read_count, gcs/request_count, gcs/request_latencies, and gcs/retry_count expose cache reuse, remote traffic, latency, and retry behavior.

After Batch accepts the job, the submission script prints the complete Batch job link and Metrics Explorer links filtered by the immutable Batch job UID. Every job receives utilization, host I/O, and cache charts. DCGM jobs also receive a GPU profiling chart and DCGM traffic series in the I/O chart. Separate charts keep percentages, byte rates, and operation rates on meaningful axes. Together they cover CPU, GPU, memory, optional GPU occupancy and execution-pipe use, network, disk, cache hit/miss, cached bytes, and downloaded bytes over the standard job’s two-week maximum window. The filters use the worker instance-name metadata retained with the time series, so deleting the temporary worker does not delete the measurements. For older jobs, change each chart’s time selector to the job interval. Agent and Cloud Storage FUSE metrics require the training service account’s roles/monitoring.metricWriter binding and the Monitoring API managed by Terraform.

Job lifecycle

The script returns after Batch accepts or rejects job creation and, after acceptance, prints the job and monitoring links. Batch creates the machine and local SSD, sends logs to Cloud Logging, and releases the temporary resources when the job succeeds, fails, or is cancelled. An operator can cancel the job from Batch → Jobs in Google Cloud Console.

The job definition does not impose a shorter task timeout. Standard-provisioned Batch jobs remain subject to the platform’s 14-day maximum running time.

The adapter and job definition are locally validated. A real cache measurement and cleanup check require a separately approved cloud run.

See Also