Command-line interface¶
RFFM currently has three source-checkout command surfaces. They have different responsibilities and are not installed console scripts.
Classification¶
Maturity: implemented
./rffm classify --artifact CLASSIFICATION_ARTIFACT \
--artifact-sha256 sha256:DIGEST --request REQUEST.json
./rffm classify --endpoint CLASSIFICATION_URL --request REQUEST.json
The request file contains one in-phase and quadrature (I/Q) window plus the
physical metadata required by the classification artifact. Local mode verifies
the caller-provided SHA-256 identity before loading the immutable classifier
artifact. Remote mode posts the same typed request to the exact endpoint. Set
RFFM_CLASSIFY_BEARER_TOKEN when that endpoint requires a standard bearer
token.
The CLI does not acquire credentials. Obtain the bearer token from the endpoint’s identity provider. For a private Cloud Run service, an authenticated principal with permission to invoke the service can set the token before calling the endpoint:
export RFFM_CLASSIFY_BEARER_TOKEN="$(gcloud auth print-identity-token)"
Success writes one classification response as JSON to standard output. The response contains the predicted label, task-ordered uncalibrated softmax scores, and exact classification-artifact and pretraining-checkpoint identities. Diagnostics and actionable input, artifact, or service errors go to standard error.
Remote mode applies the same 1 MiB (1,048,576-byte) limit to successful and error response bodies. It reads one additional byte to detect overflow and rejects an oversized response explicitly rather than parsing truncated JSON. For an oversized error response, diagnostics retain the HTTP status and reason without retaining the oversized service detail.
Bounded pretraining¶
Maturity: verified
python -m rffm.cli.main pretrain
[--config-dir PATH]
[--config-name NAME]
--experiment NAME
--run-id ID
--attempt-id ID
[--rffm-repository-root PATH]
[--gcs-bucket-namespace-root PATH]
[--source-revision SHA]
[--container-image-uri URI]
[--launch-config-uri URI]
[--launch-config-sha256 SHA256]
[--override KEY=VALUE]...
Option |
Required |
Contract |
|---|---|---|
|
no |
Hydra tree; defaults to |
|
no |
Root config name; defaults to |
|
yes |
Readable experiment config-group name |
|
yes |
Logical run identity and artifact path component |
|
yes |
One attempt within the logical run |
|
no |
Root used to resolve repository-relative paths |
|
for |
Directory whose immediate children are mounted bucket names |
provenance options |
no locally |
Source, immutable image, provider launch URI, and launch hash recorded in manifests |
|
no |
Repeatable Hydra dot-list override |
The command writes one sorted JSON result to standard output after success. Configuration, data, validation, or training failures propagate as non-zero process failures. Non-finite training failures also attempt to publish a strict diagnostic and terminal failed manifest without hiding the primary error.
Supervised post-training¶
Maturity: implemented
python -m rffm.cli.main posttrain --config POSTTRAINING.yaml
Pass --hidden-state-cache-directory SCRATCH to populate a BF16,
memory-mapped cache during the first complete training and validation passes.
Later epochs reuse those frozen-backbone states. The directory must be
attempt-local scratch storage; it is not published or referenced by terminal
evidence. After the first completed epoch, one structured log reports the
training, validation, and total cache-file sizes in bytes. Omitting the option
preserves the uncached path.
To diagnose that cache lifecycle, add the optional profiling block to a recipe that has at least two epochs and run it with the cache-directory option:
profiling:
batch_count: 10
The application writes standard PyTorch Chrome traces and a
cache_lifecycle_summary.json under application/profiling/. The bounded
traces cover the leading training and validation batches in epoch one, including
consumer wait, input preparation, the frozen backbone, cache-store calls, and
classifier work. Separate traces cover the waits that complete cache writes.
Epoch-two traces cover the first completed-cache reuse, including the
one-batch-ahead host-to-device prefetch and classifier work.
The summary complements the parent-process traces with durations measured where
the work happens: raw and completed-cache reads inside DataLoader workers, and
preallocation, staging, backpressure, CUDA-copy waits, and mmap writes inside
the cache caller and background writer. Read observations are bounded by
batch_count; cache-write observations cover the complete epoch-one cache
construction. These inclusive durations expose I/O pressure but do not break
filesystem time down into individual system calls or page faults. Profiling is
diagnostic and adds overhead, so its timings are not benchmark results. Omitting
the block preserves normal post-training behavior.
Phase durations are inclusive and non-additive: worker observations may run in parallel, and writer-task totals contain their nested wait/write phases. Use backpressure and completion waits as critical-path signals; do not sum phase totals into an epoch duration or expected runtime reduction.
The strict YAML recipe identifies a human-readable
classification-modulation-... run ID and immutable attempt ID, the canonical
RadioML2018 bundle, the exact pretraining checkpoint bytes and resolved
configuration, the existing input policy, pooling and optimizer settings,
execution identity, and the local and durable output roots. The command restores
the radio-frequency backbone through its existing next-patch-prediction task
state. It freezes that backbone, trains only pooling and the classification
head, and selects an immutable checkpoint by validation macro accuracy and then
validation loss. The recipe must declare one procedure explicitly:
procedure: validation_selection # or final_qualification
When an attempt-local cache directory is supplied, brain floating point 16-bit (BF16) is the only hidden-state cache representation. The application applies the same BF16 round trip during the first pass and restores 32-bit floating point (FP32) before classifier work in every epoch. Masks, targets, classifier parameters, gradients, optimizer state, loss, and metric reduction retain their existing data types.
Validation selection trains, selects, and publishes the artifact without constructing or iterating the test dataset. Final qualification preserves that workflow and then evaluates the reloaded selected artifact once on the untouched test split. A successful validation-selection attempt is not qualification evidence.
Provider adapters may supply the optional immutable image, provider job, and
launch-config identities through RFFM_CONTAINER_IMAGE_URI,
RFFM_PROVIDER_JOB_ID, RFFM_LAUNCH_CONFIG_URI, and
RFFM_LAUNCH_CONFIG_SHA256. Explicit YAML values take precedence.
After selection, the command packages one self-contained classification artifact. One attempt writes this canonical application layout:
<output-root>/<run-id>/attempts/<attempt-id>/application/
resolved_config.sha256_<digest>.yaml
metrics.jsonl
checkpoints/epoch_<4-digit-epoch>.pt
profiling/*.json # optional diagnostic traces and phase summary
task_artifact.pt
test_report.json # final qualification only
run_manifest.json
run_manifest.json is published last. It records the declared procedure and is
the authoritative source for the selected checkpoint, classification artifact,
datasets, source and runtime context, and every durable artifact URI, byte size,
and SHA-256. Final-qualification manifests additionally require the test dataset
and report; validation-selection manifests require both to be absent. The
command writes a small manifest-backed JSON result to standard output, with
test_accuracy set to null for validation selection. It does not submit cloud
work or deploy the selected artifact; provider adapters own those operations.
Preprocessing submission¶
Maturity: runtime-isolated
./rffm preprocess DATASET --env dev --image-uri IMAGE_URI
This root script submits the selected dataset to the configured Google Cloud
Workflow by invoking gcloud. It prints Workflow, Batch, log, and artifact
identities needed by an operator. The command requires configured Google Cloud
credentials and mutates external state; it is not part of the local
documentation quickstart.
IMAGE_URI must identify the exact published runtime image by an immutable
@sha256:... digest. Mutable tags such as :latest are not accepted.
The only current environment is dev. The dataset argument is passed to the
deployed workflow; the local parser does not claim that every registry entry is
canonicalization-ready.