Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Paper β’ 2608.13417 β’ Published β’ 51
None defined yet.
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
CAST: Game Solvers as Turn-Level Teachers for LLM Agents