search-opsd-qwen3-4b

This checkpoint documents a failed training run. It is archived for reproducibility and as a negative result, not as a usable model.

Setup

Qwen3-4B-Instruct fine-tuned with main_sdar on the SDAR-TopK search-QA task, using the "OPSD" configuration:

  • actor_rollout_ref.actor.pg_loss_coef=0 โ€” the GRPO/PG policy-gradient term is zeroed out; the actor receives no task-reward-driven gradient.
  • algorithm.sdar.gate_beta=5.0 โ€” learned per-token gating (not uniform).
  • algorithm.sdar.sdar_coef=0.01 โ€” SDAR self-distillation coefficient.
  • No top-k budget mechanism.

The policy update is driven entirely by the SDAR self-distillation term (student imitating a teacher's next-token distribution under privileged skill text), with zero grounding in whether the produced answer is actually correct.

What happened

Training ran the full 150 steps and the checkpoint is structurally intact (verified: Qwen3ForCausalLM, 36 layers, hidden 2560, vocab 151936, BF16, 4.41B params, 399 tensors). But the training run itself collapsed:

  • Steps 1-20: normal training-time episode/success_rate, peaking at 0.435 (step 20).
  • Steps 24-34: sharp decline.
  • Steps 30-35: actor/entropy_loss spikes from ~0.3-0.5 to ~10 (near-maximum entropy for the vocabulary) within 5-6 steps; response length balloons toward the 512-token cap; tool-call count climbs toward the 4-turn cap.
  • Steps 35-149 (115 steps, 77% of training): episode/success_rate is stuck at exactly 0.000 and never recovers.

Held-out evaluation (confirms the collapse, not an eval artifact)

Full held-out test set (51,713 examples, greedy decoding, 4/4 shards evaluated, checkpoint load verified rank-by-rank):

data_source n success rate tool calls/ex
popqa 14,267 0.0144 2.939
triviaqa 11,313 0.0126 2.963
nq 3,610 0.0030 2.986
2wikimultihopqa 12,576 0.0028 2.989
hotpotqa 7,405 0.0020 2.989
musique 2,417 0.0000 2.995
bamboogle 125 0.0000 2.992
micro-average 51,713 0.0079

For comparison, sibling configurations on the same task all score 0.42-0.45 micro-average success rate (baseline, RLSD, plain GRPO, DPos r=0.10, GRPO+OPSD which keeps the PG loss). Tool-call counts here sit at 2.9-3.0 across every data source, essentially pinned to the environment's 4-turn cap โ€” the model exhausts its retrieval budget on every question without ever converging on a correct answer.

Likely cause

Removing the task-reward-gradient entirely and relying solely on SDAR self-distillation appears unstable in this environment: once a training step pushes the student's output distribution far enough from the teacher's, there is no mechanism pulling the policy back toward task-correct behavior, and it does not recover. The sibling configuration that keeps the PG loss alongside the same SDAR term (gate_beta=0.0, no gating) trained stably to 0.4501 micro-average, which is the strongest contrast case for this failure mode.

Downloads last month
15
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support