Project layout¶
Maturity: implemented
This map names the current repository, not a proposed package tree.
rf-foundation-models/
common/ shared preprocessing and validation primitives
configs/
config.yaml Hydra composition root
experiment/ complete bounded experiment presets
dataset/ experiment-facing dataset selection
datasets/registry.yaml canonical dataset locations and readiness
input_pipeline/ normalization/ patching/
model/ objective/ optimizer/ training/ checkpoint/
launch/ provider launch configuration
preprocessing/ dataset-specific canonicalization adapters
src/rffm/
applications/pretrain/ bounded application and attempt lifecycle
applications/posttrain/ classifier training and checkpoint selection
cli/main.py pretraining command dispatcher
data/ canonical reader, registry, collation
inference/ task artifacts and typed classification
model/ causal model contract and tiny implementation
signal/ normalization, patching, model inputs
training/ NPP, guarded step, diagnostics, checkpoints
storage.py filesystem and mounted-GCS path resolution
types.py semantic RF sample and batch records
training/tests/ executable contracts for the bounded path
inference/ inference runtime dependency contracts
infra/runtime/ training and inference runtime definitions plus provider adapters
docs/preview/ reader-facing Sphinx source
Import boundary¶
The repository has no root pyproject.toml and does not publish an rffm
distribution. Training CI runs with PYTHONPATH=src:. and installs dependencies
from analysis/eda/pyproject.toml. The training container sets the same source
and repository paths explicitly.
Command surfaces¶
The root ./rffm script dispatches preprocessing submission and classification
from common.cli. Bounded pretraining uses the distinct Python application
entry point python -m rffm.cli.main pretrain. These commands run from a source
checkout; this repository does not install a console script.
Stable and specific homes¶
Common canonical bundle, validation, hashing, and persistence concepts live in
common/because preprocessing and training both consume them.Dataset-specific source parsing stays under
preprocessing/.Reusable learning primitives stay in the
data,signal,model, andtrainingpackages. Artifact-backed classification stays ininference.Training and service composition stay under their explicit
applicationsmodules.Gemini Enterprise Agent Platform resource and submission behavior stays under
infra/runtime.