Skip to main content
Use this page for training checkpoint and resume knobs and GRPO metric interpretation that are easy to miss when running cookbook-driven reinforcement learning. For sampling and scheduling configuration (completions per prompt, staleness budget, concurrency, loss behavior), see Cookbook: Reinforcement Learning. The canonical cookbook reference for save, resume, and promote is Checkpoints and Resume. Low-level SDK APIs are documented in Saving and loading.

dcp_save_interval

Controls how often full training state (weights and optimizer) is checkpointed using DCP (Distributed Checkpoint) format. When set to 0 (the default), no periodic DCP checkpoints are written for resume. Only sampler and HuggingFace-format weight snapshots may be produced — these preserve model weights but not optimizer state. When set to a positive integer N, a full DCP checkpoint is written every N steps. Why this matters: If a training job is interrupted, optimizer state is lost unless dcp_save_interval is set. The model resumes from the last checkpoint, but the optimizer re-initializes from scratch — which can affect training stability and effective learning rate.

Example (cookbook Config)

Some internal or forked recipes may expose the same interval on a nested config type (for example a weight sync block). The field name is always dcp_save_interval; see your recipe’s Config dataclass for the exact attribute path.

Job recovery and preemption

For transient control-plane or worker interruptions, the trainer job manager exposes reconnect_and_wait so your driver can wait for a resumable state and resume cleanly.
load_state_with_optimizer() only restores optimizer state from DCP-format checkpoints. If you point it at an HF or sampler snapshot, optimizer state silently won’t be restored. Always load from the path returned by save_state() when you need full optimizer restore. See Saving and loading.

Metrics reference

ppo_kl vs ref_kld

GRPO training logs two KL divergence metrics that measure different things: Which one to monitor: ref_kld is the metric to watch for policy drift. A sudden large jump in ref_kld may indicate reward hacking or that the KL penalty coefficient needs tuning. The cookbook does not always surface ref_kld by default. To add it, you can use the k3 unbiased estimator: