HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 2 days ago • 115
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 5 days ago • 269
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models Paper • 2608.10538 • Published 8 days ago • 14
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation Paper • 2608.12990 • Published 6 days ago • 14
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published 6 days ago • 49
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence Paper • 2608.12743 • Published 6 days ago • 42
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published 19 days ago • 105
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Paper • 2608.11079 • Published 8 days ago • 16
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory Paper • 2608.07169 • Published 12 days ago • 49
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published 13 days ago • 95
FinanceHarness: Autonomous Financial Deep Research Framework Paper • 2607.27853 • Published 12 days ago • 11
FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory Paper • 2608.04530 • Published 14 days ago • 14
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents Paper • 2608.04574 • Published 14 days ago • 16