-
Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
Paper • 2603.04597 • Published • 211 -
SII-Enigma/Llama3.2-8B-Ins-AMPO
Text Generation • 8B • Updated • 8 -
Understanding R1-Zero-Like Training: A Critical Perspective
Paper • 2503.20783 • Published • 60 -
Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
Paper • 2509.25779 • Published • 19
ming
elonming
AI & ML interests
None yet
Organizations
None yet
DailyPapers
-
Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning
Paper • 2603.04597 • Published • 211 -
SII-Enigma/Llama3.2-8B-Ins-AMPO
Text Generation • 8B • Updated • 8 -
Understanding R1-Zero-Like Training: A Critical Perspective
Paper • 2503.20783 • Published • 60 -
Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
Paper • 2509.25779 • Published • 19
models 0
None public yet
datasets 0
None public yet