Logging and observability¶
Maturity: verified for application records; runtime-isolated for provider log archives
The bounded application publishes structured files and emits concise runtime progress through Python logging. It does not provide a general experiment tracker. The command-line entry point sends logs to standard error, leaving its final machine-readable result on standard output. A Vertex CustomJob captures that standard-error stream in Cloud Logging for live monitoring.
File |
Publication |
Meaning |
|---|---|---|
|
before training |
complete validated application recipe |
|
terminal, atomic |
completed optimizer-step records only |
|
at cadence/final step |
model, optimizer, and lineage state |
|
non-finite failure |
exact failure stage and tensor diagnostics |
|
terminal, last |
status, identities, lineage, and artifact hashes |
|
after failed CustomJob |
compressed Cloud Logging entries for the exact job |
|
after log archive |
failed-manifest, job, query, and archive identities |
|
after completed correction work |
human-readable evidence, decision, implementation, and validation lineage |
Step records¶
Each JSON Lines record contains:
{
"event_type": "optimizer_step",
"global_step": 1,
"loss": 0.125,
"num_target_patches": 8,
"sample_ids": ["sample-001"]
}
Serialization rejects non-standard non-finite JSON values. Component metrics
cannot redefine application-owned event_type, global_step, or loss
fields.
Terminal status¶
A successful manifest records status="succeeded" and final global_step. A
non-finite failure records status="failed", the last completed step, failure
stage, and attempted step. The manifest is the terminal inventory.
Runtime logs report attempt start and completion, every completed optimizer step, validation start and completion, periodic validation progress, checkpoint publication, and terminal-manifest publication. Validation progress includes completed batches, completed RF samples, percentage, elapsed time, estimated remaining time, and validation samples per second. Each completed training-step log includes end-to-end step duration, time waiting for its input batch, and training samples per second. A completed checkpoint-publication log includes publication duration. These executor-independent fields are operational telemetry rather than durable training evidence; metrics, checkpoints, and the terminal manifest remain the retained application record.
Provider logs remain provider-owned during execution. After a failed CustomJob, the separate evidence command can archive that exact job’s Cloud Logging output without changing the submission adapter or terminal application manifest. Training code does not parse or execute the later Markdown correction record.