Architecture

Maturity: verified

RFFM separates reusable learning primitives from one executable pretraining workflow and from provider submission. That boundary keeps a canonical dataset, model, objective, or checkpoint usable without importing application lifecycle or Google Cloud behavior.

Pretraining data flow

Validated pretraining application and provider data flow

For managed execution, an independent shell adapter preserves the launch configuration and submits the immutable training image as one Gemini Enterprise Agent Platform CustomJob before handing control to python -m rffm.cli.main pretrain.

The adapter supplies provider identities and the mounted Cloud Storage bucket namespace. It does not calculate loss, decide checkpoint compatibility, or write application metrics.

Managed failure and correction flow

The managed failure path crosses ownership boundaries without turning correction into training automation:

Managed failure and correction ownership architecture

The failed terminal manifest is the immutable handoff from the application to provider evidence collection. The completed Markdown record is the immutable handoff from correction work to future humans and agents. It is not an input to training and does not approve or launch another attempt.

The storage layout shows the files produced by each owner. The repository README gives the operator commands, and the failed-attempt evidence and correction record defines the detailed evidence contract.

Responsibility layers

Canonical data and signal preparation

src/rffm/data/ reads the common canonical bundle and returns semantic RF records. src/rffm/signal/ owns I/Q normalization, physical-metadata normalization, fixed patching, and the model-input record. These modules do not select an experiment or write run artifacts.

Model and training primitives

src/rffm/model/ owns the causal-patch model contract and the shipped tiny implementation. src/rffm/training/ owns next-patch prediction, guarded optimizer steps, bounded diagnostics, and generic checkpoint save/load.

These are reusable library capabilities. They do not own run IDs, provider job IDs, application directories, or terminal manifests.

Pretraining application

src/rffm/applications/pretrain/ composes configuration, selects code-owned component registrations, creates datasets and dataloaders, owns run and attempt identity, executes the bounded loop, publishes artifacts atomically, and validates resume compatibility before state mutation.

The command dispatcher in src/rffm/cli/main.py is an application entry point, not a general-purpose training library API.

Provider adapter

The submission scripts validate separate launch YAML files, require an immutable image digest and full source revision, and persist the selected launch configuration. submit_cloud_batch_training.sh submits configured single-GPU work with a dataset cache; submit_vertex_ai_training.sh keeps the existing CustomJob path. An operator requests cancellation; the selected provider stops the job and removes its temporary worker and disks.

Capability limitations

There is no landed train/validation epoch loop, scheduler, latest/best checkpoint selection, exact sampler or random-number-generator resume, dataset mixture, distributed strategy, sweep runner, callback system, post-training application, or evaluator. Those concepts remain in the design sources and are not part of the current application graph.

See Also