Gemini Enterprise Agent Platform serverless training¶
Maturity: verified for the bounded submission and recorded one-step/resume evidence.
Gemini Enterprise Agent Platform (formerly Vertex AI) is the Google Cloud offering used for the repository’s serverless training submission path.
Launch configurations¶
infra/runtime/submit_vertex_ai_training.sh is the committed provider adapter.
It reads the file selected by LAUNCH_CONFIG_PATH, whose provider value is
vertex_ai, uploads that exact file, and calls
gcloud ai custom-jobs create to create a CustomJob. This command mutates
cloud state and is for an authorized operator, not the local getting-started
path.
Launch selection¶
The committed configurations select one Gemini Enterprise Agent Platform
CustomJob replica in rf-foundation-models-dev / us-central1 using
rf-fm-training-dev@rf-foundation-models-dev.iam.gserviceaccount.com.
Configuration |
Machine |
Accelerator |
Intended use |
|---|---|---|---|
|
|
one |
routine Small training and evaluation |
|
|
one |
Small training when standard L4 capacity is unavailable; waits up to one day |
|
|
one |
bfloat16-capable capacity fallback or larger workloads |
Required submission inputs¶
The script requires DISPLAY_NAME, RUN_ID, ATTEMPT_ID, EXPERIMENT,
IMAGE_URI, SOURCE_REVISION, LAUNCH_CONFIG_PATH, and
LAUNCH_CONFIG_URI. IMAGE_URI must end in an immutable 64-hex SHA-256 digest;
SOURCE_REVISION must be a full 40-hex Git commit.
For resume, set both RESUME_FROM and positive TARGET_MAX_STEPS. Omitting
both selects a fresh attempt. Setting only one is rejected.
To change settings in the selected experiment, set CONFIG_OVERRIDES_PATH to
a text file containing one non-empty key=value override per line. The adapter
passes each line through the application’s existing --override option. The
application validates the resulting configuration before training and stores
the complete resolved configuration and its hash with the run artifacts.
Do not repeat training.max_steps or checkpoint.resume_from in this file
when resuming. Use TARGET_MAX_STEPS and RESUME_FROM for those two settings.
Submission behavior¶
Before submission, the script validates the exact launch fields, identifiers,
image and source immutability, accelerator and replica counts, optional
setting file, and resume pair.
It uploads the launch file with --if-generation-match=0, computes its SHA-256,
and passes source, image, launch, run, attempt, experiment, and mounted /gcs
inputs to the application.
The Flex Start configuration adds Vertex’s FLEX_START scheduling strategy and
a 24-hour maximum wait for L4 capacity. Waiting and provisioning remain
provider-owned; the submission adapter still returns immediately and does not
retry or relaunch jobs.
The script returns after the provider accepts or rejects job creation. It does not wait, monitor, retry, cancel, verify terminal artifacts, or clean up the job. Those actions remain operator responsibilities.
Bounded execution evidence¶
Recorded bounded execution shows one fresh optimizer step and a separate attempt resumed to step two on a Gemini Enterprise Agent Platform L4 worker. It validates the bounded path; it does not establish train/validation, model-quality, exact-replay, or scale claims.