TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Paper • 2608.20958 • Published 7 days ago • 58
CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing Paper • 2608.17566 • Published 10 days ago • 15
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation Paper • 2604.03738 • Published Apr 4 • 1
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation Paper • 2510.18692 • Published Oct 21, 2025 • 41