The worker dispatches on the recipe’s schema version and model.implementation. Keep both identifiers with the recipe and exported artifacts.
Strategy versions
| Implementation | Computation version |
|---|---|
lfm2-full-weight-two-stage-v1 |
7.0.0 through 11.0.0 |
lfm2-packed-sft-v1 |
12.0.0 |
lfm2-packed-qat-v1 |
13.0.0 |
Use minifield-worker --list-strategies to inspect the installed registry. Direct worker recipes select any installed strategy; the supervisor advertises its accepted computations when claiming work.
Two-stage full-weight training
Two-stage training learns action routing and argument generation with separate supervision. It updates the model body while keeping the tied embedding and language-model head frozen.
training.updates sets the update count; set limits.max_steps to the same value. sequence_length defaults to 8192, with a supported range of 2–32768. The default training seed is 42.
The dense path uses quantization_profile: "disabled-v1" and schedule: "constant-v1". Forward computation uses BF16; master weights and optimizer state use FP32.
Packed SFT settings
Packed SFT trains complete action targets and updates every parameter leaf, including the tied embedding. It uses recipe computation 12.0.0 and the same dense numerical profile.
The training object contains the physical layout and numerical settings:
| Field | Type | Default |
|---|---|---|
compute_dtype |
"bfloat16" |
"bfloat16" |
gradient_accumulation_steps |
integer (> 0) | Required |
master_dtype |
"float32" |
"float32" |
max_projection_chunks |
integer (> 0) | Required |
optimizer |
object (Optimizer) | Optional |
projection_chunk_size |
integer (> 0) | Required |
quantization_profile |
"disabled-v1" |
"disabled-v1" |
rows_per_microbatch |
integer (> 0) | Required |
schedule |
"constant-v1" |
"constant-v1" |
seed |
integer (≥ 0, ≤ 18446744073709551615) | 42 |
sequence_length |
integer (≥ 8, ≤ 32768) | 8192 |
Each row holds sequence_length tokens. rows_per_microbatch sets the row count, and gradient_accumulation_steps sets microbatches per optimizer update.
projection_chunk_size × max_projection_chunks sets supervised target capacity per microbatch. seed controls the training shuffle.
For example, this training fragment uses 1 row per microbatch, accumulates 4 microbatches per update, and allows 2,048 supervised target positions per microbatch:
{
"sequence_length": 8192,
"rows_per_microbatch": 1,
"gradient_accumulation_steps": 4,
"projection_chunk_size": 256,
"max_projection_chunks": 8,
"seed": 42,
"compute_dtype": "bfloat16",
"master_dtype": "float32",
"quantization_profile": "disabled-v1",
"schedule": "constant-v1"
}Choose dimensions against your model, example lengths, and GPU memory. generation.json records the prepared packing plan; training events report its input and supervised token counts.
Packed quantization-aware training
Packed QAT uses the direct worker’s 13.0.0 recipe and the packed physical settings above. Set quantization_profile to ternary-g128-v1 and schedule to immediate-v1.
Eligible body weights participate in the forward pass as ternary matrices with group-128 scales. FP32 master weights receive the optimizer updates. The tied embedding stays frozen and exports as BF16.
The objective combines supervised cross-entropy with distillation from the job’s dense starting weights. Its recipe settings are:
| Field | Type | Default |
|---|---|---|
distill_weight |
number (≥ 0) | 1 |
distill_temperature |
number (> 0) | 2 |
ce_weight |
number (≥ 0) | 0.1 |
scale_rule |
"absmax" or "absmean" or null |
null |
A null scale rule uses the quantization profile’s rule. A zero objective weight disables that term. These settings participate in the computation identity.
The exported body uses 2-bit packed ternary codes and FP16 scales in the minifield.ternary.v1 format. Preserve the quantization settings with the export identity when comparing it in Runtime.
Optimizer defaults
The training.optimizer object uses AdamW. An omitted optimizer uses these defaults:
| Field | Type | Default |
|---|---|---|
beta1 |
number (≥ 0, < 1) | 0.9 |
beta2 |
number (≥ 0, < 1) | 0.95 |
clip |
number (> 0) | 1 |
eps |
number (> 0) | 1e-8 |
learning_rate |
number (> 0) | 0.00001 |
weight_decay |
number (≥ 0) | 0.01 |
Training records gradient norm, update norm, loss, token counts, and elapsed step time. Use these alongside held-out evaluation to judge a run.
Limits and recovery
Packed recipes checkpoint every 100 optimizer updates by default. Set limits.checkpoint_every_steps to choose the cadence and limits.max_steps to set an optional update cap.
| Field | Type | Default |
|---|---|---|
checkpoint_every_steps |
integer (> 0) | 100 |
max_steps |
integer (> 0) or null | null |
The computation identity binds model and data references, serialization, packing, numerical settings, optimizer, schedule, and training seed. Resume using the same computation settings and a new attempt identity. A changed computation starts a new run.
Read checkpoint recovery for the saved state and verification process.