Allspark: Weak to Strong Transfer via Alternating Chain of Thought Paper • 2609.32913 • Published 6 days ago • 4
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Paper • 2608.27454 • Published Aug 27 • 36
Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow Paper • 2209.03003 • Published Sep 7, 2022 • 4
CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasks Paper • 2606.21777 • Published Jun 19 • 5
No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions Paper • 2606.13044 • Published Jun 11 • 11
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning Paper • 2603.12529 • Published Mar 13 • 19
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset Paper • 2403.12945 • Published Mar 19, 2024 • 2
Mini-BEHAVIOR: A Procedurally Generated Benchmark for Long-horizon Decision-Making in Embodied AI Paper • 2310.01824 • Published Oct 3, 2023 • 1
Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning Paper • 2603.11653 • Published Mar 12 • 2
EntRGi: Entropy Aware Reward Guidance for Diffusion Language Models Paper • 2602.05000 • Published Feb 4 • 2