Models and objectives¶
Maturity: verified
RFFM keeps the model boundary separate from the training objective and optimizer step. A model transforms prepared inputs into outputs, an objective turns those outputs into a loss, and an optimizer step updates parameters while the caller retains ownership of the run lifecycle.
Model implementations¶
The application ships a mainline causal I/Q model and retains a Tiny model as a fast regression fixture. Neither model includes a tokenizer or pretrained language weights.
Model input and sequence¶
Both implementations use CausalPatchModel as their model contract.
Its prepared sequence contains three masked physical-metadata positions followed
by fixed I/Q patches:
[sample rate] [center frequency] [bandwidth] [patch 0] ... [patch N-1]
Given metadata (B, M) and patches (B, N, 2P), CausalPatchModel returns
hidden states (B, M + N, D), the metadata prefix length, and a boolean mask
aligned to the sequence. State i may depend only on valid positions through
i.
The qwen3_causal_iq model projects physical metadata and I/Q patches to a
shared token width, then passes those tokens through a public dense Qwen3
decoder backbone. The backbone contextualizes tokens with causal attention; it
does not map hidden states back to physical I/Q values. The separate
next-patch-prediction task owns that prediction head and loss.
The Small preset uses 16 decoder layers, width 384, six query heads, two shared key/value heads, head width 64, and SwiGLU width 1024. It provides pre-RMSNorm, per-head QK-Norm, RoPE, grouped-query causal attention, zero dropout, and bias-free transformer projections. Its roughly 25 million parameters start from random initialization.
The tiny_causal_iq regression fixture constructs:
one scalar-to-token projection for metadata;
one
2P-to-token projection for I/Q patches;learned absolute position embeddings;
a PyTorch transformer encoder driven by a causal mask;
GELU, a feed-forward width of
2D, zero dropout, and the configured number of layers and attention heads.
The Tiny preset uses P=64, D=64, two layers, four heads, and a maximum
combined sequence length of 64. These are test-path choices, not a mainline
model recommendation.
Next-patch prediction¶
NextPatchPredictionTask uses patch hidden state i to predict input patch
i+1. Its compact GELU head maps D -> max(1, D/2) -> 2P.
Each target patch is normalized over valid I/Q elements only. The mean-squared error is averaged only where the source patch, target patch, and target element are valid. Partial-patch padding therefore contributes neither to target normalization nor loss.
Optimizer step¶
The shipped adamw registration constructs PyTorch AdamW with YAML-owned
learning rate and weight decay. The reusable step checks finite input tensors,
pre-step parameters, forward outputs and loss, gradients, and post-update
parameters. A completed step returns scalar loss and valid target-patch count.
The caller owns iteration, global step, metrics, checkpoints, and failure artifacts.