JumpStart Your Policy Learning with Lessons from 160,000 Training Runs
Abstract
Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has shown that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties, but these sources of variability have not been systematically investigated together at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 160,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance among the strongest methods is often close, but the leaders differ substantially across environments. We find that proper hyperparameter tuning frequently reshuffles perceived algorithm rankings and that benchmark composition can produce conflicting conclusions. We also study hyperparameter sensitivity and transfer across environments, identifying a simple strategy for deriving strong default configurations. We use our findings to develop a dataset-conditioned recommender that provides task-specific algorithm recommendations for practitioners. Finally, we release JumpStart: a resource suite containing every trained policy, per-model scores and hyperparameters, strong baselines across all environments, training and evaluation code, and an extensible website for retrieving, analyzing, and contributing results. Together, these resources aim to make offline policy-learning research more reliable and enable future work beyond the scope of this study.
Community
We release over 160,000 models trained with offline RL across a wide range of tasks!
Prior work has shown that results in offline policy learning (e.g., robotics) are sensitive to reporting choices, hyper-parameter tuning, and properties of the dataset. However, these sources have not been systematically evaluated at the scale needed to understand how they shape the conclusions of a paper.
To address this, we train over 160,000 models across 114 diverse datasets/environments, and present an in-depth study of the factors that lead to reliable evaluation and strong model performance. We make all models available for public use for the community to build on.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics (2026)
- DataFlex-RL: An Evaluation Platform for RLVR Data Policies (2026)
- Training seeds and model-selection stability in recommender-system evaluation (2026)
- Bandits in Prod: Hyperparameter Optimization at Inference Time (2026)
- EasyPPO: Stabilizing the Critic Is Key (2026)
- Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning (2026)
- On BatchNorm Forward Modes in Value-Based Reinforcement Learning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.13730 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper