Gemini Enterprise Agent Platform serverless training

Maturity: verified for the bounded submission and recorded one-step/resume evidence.

Gemini Enterprise Agent Platform (formerly Vertex AI) is the Google Cloud offering used for the repository’s serverless training submission path.

Launch configurations

infra/runtime/submit_vertex_ai_training.sh is the committed provider adapter. It reads the file selected by LAUNCH_CONFIG_PATH, whose provider value is vertex_ai, uploads that exact file, and calls gcloud ai custom-jobs create to create a CustomJob. This command mutates cloud state and is for an authorized operator, not the local getting-started path.

Launch selection

The committed configurations select one Gemini Enterprise Agent Platform CustomJob replica in rf-foundation-models-dev / us-central1 using rf-fm-training-dev@rf-foundation-models-dev.iam.gserviceaccount.com.

Configuration

Machine

Accelerator

Intended use

configs/launch/vertex_ai_l4.yaml

g2-standard-4

one NVIDIA_L4

routine Small training and evaluation

configs/launch/vertex_ai_l4_flex_start.yaml

g2-standard-4

one NVIDIA_L4

Small training when standard L4 capacity is unavailable; waits up to one day

configs/launch/vertex_ai_a100.yaml

a2-highgpu-1g

one NVIDIA_TESLA_A100

bfloat16-capable capacity fallback or larger workloads

Required submission inputs

The script requires DISPLAY_NAME, RUN_ID, ATTEMPT_ID, EXPERIMENT, IMAGE_URI, SOURCE_REVISION, LAUNCH_CONFIG_PATH, and LAUNCH_CONFIG_URI. IMAGE_URI must end in an immutable 64-hex SHA-256 digest; SOURCE_REVISION must be a full 40-hex Git commit.

For resume, set both RESUME_FROM and positive TARGET_MAX_STEPS. Omitting both selects a fresh attempt. Setting only one is rejected.

To change settings in the selected experiment, set CONFIG_OVERRIDES_PATH to a text file containing one non-empty key=value override per line. The adapter passes each line through the application’s existing --override option. The application validates the resulting configuration before training and stores the complete resolved configuration and its hash with the run artifacts.

Do not repeat training.max_steps or checkpoint.resume_from in this file when resuming. Use TARGET_MAX_STEPS and RESUME_FROM for those two settings.

Submission behavior

Before submission, the script validates the exact launch fields, identifiers, image and source immutability, accelerator and replica counts, optional setting file, and resume pair. It uploads the launch file with --if-generation-match=0, computes its SHA-256, and passes source, image, launch, run, attempt, experiment, and mounted /gcs inputs to the application.

The Flex Start configuration adds Vertex’s FLEX_START scheduling strategy and a 24-hour maximum wait for L4 capacity. Waiting and provisioning remain provider-owned; the submission adapter still returns immediately and does not retry or relaunch jobs.

The script returns after the provider accepts or rejects job creation. It does not wait, monitor, retry, cancel, verify terminal artifacts, or clean up the job. Those actions remain operator responsibilities.

Bounded execution evidence

Recorded bounded execution shows one fresh optimizer step and a separate attempt resumed to step two on a Gemini Enterprise Agent Platform L4 worker. It validates the bounded path; it does not establish train/validation, model-quality, exact-replay, or scale claims.

See Also