Browse docs
DocumentationTraining

Evaluation

Measure held-out loss, decoded actions, rejection behavior, and training usage.

Evaluate on case families held out before wording expansion. Keep the action contract and serialization identical to the training recipe, then compare the exact exported model in Runtime.

Configure packed evaluation

Packed recipes separate routine decoding from final decoding. For example, this evaluation fragment evaluates every 100 updates, decodes 3 routine samples, and decodes 20 at the end:

packed-evaluation.json
{
  "every_steps": 100,
  "samples": 20,
  "max_new_tokens": 256,
  "routine": {
    "samples": 3,
    "max_new_tokens": 128
  }
}

Routine sample and token budgets must fit within their final counterparts. A samples value of 0 runs objective evaluation with zero qualitative decodes.

Field Type Default
every_steps integer (> 0) Required
max_new_tokens integer (> 0) Required
routine object (RoutineEvaluation) Required
samples integer (≥ 0) 3
routine.max_new_tokens integer (> 0) 256
routine.samples integer (≥ 0) 1

Read each phase

Phase Work performed
Baseline Full held-out objective at step 0
Routine Full held-out objective plus the routine decode budget
Final Full held-out objective plus the final decode budget

Baseline evaluation uses zero qualitative decodes. The worker writes phase results under evaluations/, and copies the final result to evaluation.json and the model export.

Packed objective loss is mean negative log-likelihood over all supervised holdout tokens. The report includes supervised_tokens, input_tokens, packed_rows, physical_slots, padding_tokens, and packing_density so the denominator and packing remain visible.

Inspect decoded actions

Qualitative evaluation greedily decodes an action object and its assistant end token. It checks the output’s structure, action name, and arguments against the expected decision.

Metric Interpretation
qualitative_samples Number of sampled decisions decoded
action_successes Decodes with the expected action and arguments
action_success_rate Successful actions as a fraction of decoded samples
false_rejections A control action emitted when a non-control action was expected
false_acceptances A known non-control action emitted when a control action was expected
qualitative_decoded_tokens Tokens generated during sampled evaluation

Failure categories distinguish wrong_route, wrong_arguments, invalid_json, non_object, missing_fields, extra_fields, empty, and unterminated. Inspect the individual decodes as well as the aggregate categories.

Each decode records its expected and emitted route, expected and emitted kind, parsed action, raw generated text, terminated flag, and token count.

Action success requires the exact expected route and argument object, plus the assistant end token. A correct JSON object that exhausts the token budget before emitting that end token is unterminated.

The qualitative sample consists of the first configured number of holdout decisions, in their fixed order. Set samples to cover the full holdout when you need every decision decoded. With zero qualitative samples, action_success_rate is null, and objective loss still covers all supervised holdout tokens.

Keep clarification, rejection, and permission outcomes distinct in the expected actions. Exact action matching then exposes which boundary the model crossed.

Evaluate two-stage training

Two-stage reports keep routing and argument objectives separate. Routing metrics include route_pair_accuracy and route_set_exact_accuracy; argument evaluation includes teacher_forced_argument_token_accuracy.

Qualitative metrics such as qualitative_route_exact_accuracy measure decoded behavior. Compare those decoded results alongside objective metrics and inspect the saved examples before selecting an export.

Account for work

Training usage records assistant target tokens, non-padding input tokens, optimizer steps, and prepared records. Evaluation usage tracks its own forward calls and decoding work.

A resumed attempt reports the work performed during that attempt. Its inherited checkpoint identifies the earlier progress separately. Keep both attempt records when calculating the total work for a run.

Qualify the exported model

Verify the exported artifacts, then run held-out cases through the target Runtime backend. Use the exported tokenizer and prompt serializer with the same public context and action catalog.

Measure task success, rejection, latency, and memory on the device and weight format you plan to deliver. Keep the model hashes with those results so a release can point to the exact bytes evaluated.

Search documentation

Type to search the documentation.