UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
Paper • 2609.38721 • Published • 274
None defined yet.
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets