SWE-bench Collection SWE-bench (Lite, Verified, Multimodal, Multilingual) all in one place! • 5 items • Updated Dec 14, 2025 • 16
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • 4 days ago • 105
view article Article LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge LiquidAI • 5 days ago • 42
Muse Glimmer Collection Muse Glimmer 30B: multimodal agentic model for local deployment. BF16 weights, GGUF k-quants, ExecuTorch builds, DFlash drafter. • 4 items • Updated 7 days ago • 97
view article Article Meta is back with Muse Glimmer: local, agentic, multimodal, and open source +2 pcuenq, merve, burtenshaw, ariG23498 • 8 days ago • 102
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? Paper • 2605.11086 • Published May 11 • 3
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • 22 days ago • 473
view article Article LFM2.5-Encoders for Fast Long-Context Inference on CPU LiquidAI • 20 days ago • 68
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 65
view article Article ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration ibm-research • Jun 30 • 26
view article Article Featuring Every Eval Ever Results on Hugging Face Model Pages +5 deepmage121, evijit, SaylorTwift, janbatzner, borgr, irenesolaiman, julien-c • Jun 30 • 52