Tables as Images? Exploring the Strengths and Limitations of LLMs on Multimodal Representations of Tabular Data Paper • 2402.12424 • Published Feb 19, 2024
The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs Paper • 2606.18656 • Published Jun 17
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published Aug 8 • 30
Graph Pre-training for AMR Parsing and Generation Paper • 2203.07836 • Published Mar 15, 2022 • 1
Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect Paper • 2208.10099 • Published Aug 22, 2022
HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models Paper • 2310.16755 • Published Oct 25, 2023
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 22 days ago • 78
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning Paper • 2605.06326 • Published May 7 • 26
P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads Paper • 2602.09443 • Published Feb 10 • 59