RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 10 days ago • 279
Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features Paper • 2609.33463 • Published 14 days ago • 12
Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers Paper • 2608.26762 • Published 15 days ago • 23
Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models Paper • 2609.39820 • Published 11 days ago • 30
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 15 days ago • 326
Follow the Entities: A Corpus Map for Agentic Search Paper • 2609.37226 • Published 12 days ago • 100