-
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265 -
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
Paper • 2505.20139 • Published • 19 -
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
Paper • 2605.27922 • Published -
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26
Collections
Discover the best community collections!
Collections including paper arxiv:2604.08523
-
Recursive Multi-Agent Systems
Paper • 2604.25917 • Published • 289 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 91 -
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
Paper • 2604.02721 • Published • 640 -
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265
-
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265 -
Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents
Paper • 2604.06132 • Published • 122 -
FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios
Paper • 2604.07413 • Published • 98 -
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
Paper • 2604.02648 • Published • 48
-
BitNet: Scaling 1-bit Transformers for Large Language Models
Paper • 2310.11453 • Published • 108 -
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Paper • 2310.11511 • Published • 79 -
In-Context Learning Creates Task Vectors
Paper • 2310.15916 • Published • 43 -
Matryoshka Diffusion Models
Paper • 2310.15111 • Published • 46
-
From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company
Paper • 2604.22446 • Published • 125 -
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Paper • 2604.23781 • Published • 33 -
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Paper • 2604.08377 • Published • 294 -
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265
-
WildDet3D: Scaling Promptable 3D Detection in the Wild
Paper • 2604.08626 • Published • 248 -
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
Paper • 2604.06870 • Published • 44 -
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265
-
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265 -
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
Paper • 2505.20139 • Published • 19 -
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
Paper • 2605.27922 • Published -
AgentBench: Evaluating LLMs as Agents
Paper • 2308.03688 • Published • 26
-
Recursive Multi-Agent Systems
Paper • 2604.25917 • Published • 289 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 91 -
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
Paper • 2604.02721 • Published • 640 -
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265
-
From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company
Paper • 2604.22446 • Published • 125 -
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Paper • 2604.23781 • Published • 33 -
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Paper • 2604.08377 • Published • 294 -
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265
-
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265 -
Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents
Paper • 2604.06132 • Published • 122 -
FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios
Paper • 2604.07413 • Published • 98 -
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
Paper • 2604.02648 • Published • 48
-
WildDet3D: Scaling Promptable 3D Detection in the Wild
Paper • 2604.08626 • Published • 248 -
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
Paper • 2604.06870 • Published • 44 -
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Paper • 2604.08523 • Published • 265
-
BitNet: Scaling 1-bit Transformers for Large Language Models
Paper • 2310.11453 • Published • 108 -
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Paper • 2310.11511 • Published • 79 -
In-Context Learning Creates Task Vectors
Paper • 2310.15916 • Published • 43 -
Matryoshka Diffusion Models
Paper • 2310.15111 • Published • 46