Architecture¶
Maturity: verified
RFFM separates reusable learning primitives from one executable pretraining workflow and from provider submission. That boundary keeps a canonical dataset, model, objective, or checkpoint usable without importing application lifecycle or Google Cloud behavior.
Pretraining data flow¶
For managed execution, an independent shell adapter preserves the launch
configuration and submits the immutable training image as one Gemini Enterprise
Agent Platform CustomJob before handing control to
python -m rffm.cli.main pretrain.
The adapter supplies provider identities and the mounted Cloud Storage bucket namespace. It does not calculate loss, decide checkpoint compatibility, or write application metrics.
Managed failure and correction flow¶
The managed failure path crosses ownership boundaries without turning correction into training automation:
The failed terminal manifest is the immutable handoff from the application to provider evidence collection. The completed Markdown record is the immutable handoff from correction work to future humans and agents. It is not an input to training and does not approve or launch another attempt.
The storage layout shows the files produced by each owner. The repository README gives the operator commands, and the failed-attempt evidence and correction record defines the detailed evidence contract.
Responsibility layers¶
Canonical data and signal preparation¶
src/rffm/data/ reads the common canonical bundle and returns semantic RF
records. src/rffm/signal/ owns I/Q normalization, physical-metadata
normalization, fixed patching, and the model-input record. These modules do not
select an experiment or write run artifacts.
Model and training primitives¶
src/rffm/model/ owns the causal-patch model contract and the shipped tiny
implementation. src/rffm/training/ owns next-patch prediction, guarded
optimizer steps, bounded diagnostics, and generic checkpoint save/load.
These are reusable library capabilities. They do not own run IDs, provider job IDs, application directories, or terminal manifests.
Pretraining application¶
src/rffm/applications/pretrain/ composes configuration, selects code-owned
component registrations, creates datasets and dataloaders, owns run and attempt
identity, executes the bounded loop, publishes artifacts atomically, and
validates resume compatibility before state mutation.
The command dispatcher in src/rffm/cli/main.py is an application entry point,
not a general-purpose training library API.
Provider adapter¶
The submission scripts validate separate launch YAML files, require an immutable
image digest and full source revision, and persist the selected launch
configuration. submit_cloud_batch_training.sh submits configured single-GPU
work with a dataset cache; submit_vertex_ai_training.sh keeps the existing CustomJob
path. An operator requests cancellation; the selected provider stops the job
and removes its temporary worker and disks.
Capability limitations¶
There is no landed train/validation epoch loop, scheduler, latest/best checkpoint selection, exact sampler or random-number-generator resume, dataset mixture, distributed strategy, sweep runner, callback system, post-training application, or evaluator. Those concepts remain in the design sources and are not part of the current application graph.