Models and objectives

Maturity: verified

RFFM keeps the model boundary separate from the training objective and optimizer step. A model transforms prepared inputs into outputs, an objective turns those outputs into a loss, and an optimizer step updates parameters while the caller retains ownership of the run lifecycle.

Model implementations

The application ships a mainline causal I/Q model and retains a Tiny model as a fast regression fixture. Neither model includes a tokenizer or pretrained language weights.

Model input and sequence

Both implementations use CausalPatchModel as their model contract. Its prepared sequence contains three masked physical-metadata positions followed by fixed I/Q patches:

[sample rate] [center frequency] [bandwidth] [patch 0] ... [patch N-1]

Given metadata (B, M) and patches (B, N, 2P), CausalPatchModel returns hidden states (B, M + N, D), the metadata prefix length, and a boolean mask aligned to the sequence. State i may depend only on valid positions through i.

The qwen3_causal_iq model projects physical metadata and I/Q patches to a shared token width, then passes those tokens through a public dense Qwen3 decoder backbone. The backbone contextualizes tokens with causal attention; it does not map hidden states back to physical I/Q values. The separate next-patch-prediction task owns that prediction head and loss.

The Small preset uses 16 decoder layers, width 384, six query heads, two shared key/value heads, head width 64, and SwiGLU width 1024. It provides pre-RMSNorm, per-head QK-Norm, RoPE, grouped-query causal attention, zero dropout, and bias-free transformer projections. Its roughly 25 million parameters start from random initialization.

The tiny_causal_iq regression fixture constructs:

  • one scalar-to-token projection for metadata;

  • one 2P-to-token projection for I/Q patches;

  • learned absolute position embeddings;

  • a PyTorch transformer encoder driven by a causal mask;

  • GELU, a feed-forward width of 2D, zero dropout, and the configured number of layers and attention heads.

The Tiny preset uses P=64, D=64, two layers, four heads, and a maximum combined sequence length of 64. These are test-path choices, not a mainline model recommendation.

Next-patch prediction

NextPatchPredictionTask uses patch hidden state i to predict input patch i+1. Its compact GELU head maps D -> max(1, D/2) -> 2P.

Each target patch is normalized over valid I/Q elements only. The mean-squared error is averaged only where the source patch, target patch, and target element are valid. Partial-patch padding therefore contributes neither to target normalization nor loss.

Optimizer step

The shipped adamw registration constructs PyTorch AdamW with YAML-owned learning rate and weight decay. The reusable step checks finite input tensors, pre-step parameters, forward outputs and loss, gradients, and post-update parameters. A completed step returns scalar loss and valid target-patch count.

The caller owns iteration, global step, metrics, checkpoints, and failure artifacts.

See Also