Collections
Discover the best community collections!
Collections including paper arxiv:2609.02886
-
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Paper • 2608.26103 • Published • 26 -
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
Paper • 2608.27550 • Published • 95 -
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Paper • 2608.27549 • Published • 53 -
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
Paper • 2608.25518 • Published • 196
-
Wolf: Captioning Everything with a World Summarization Framework
Paper • 2407.18908 • Published • 33 -
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Paper • 2407.19985 • Published • 37 -
TPDiff: Temporal Pyramid Video Diffusion Model
Paper • 2503.09566 • Published • 45 -
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
Paper • 2506.07464 • Published • 14
-
VibeVoice Technical Report
Paper • 2508.19205 • Published • 179 -
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
Paper • 2509.22186 • Published • 177 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 199 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 779
-
WorldVLA: Towards Autoregressive Action World Model
Paper • 2506.21539 • Published • 40 -
LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
Paper • 2509.05263 • Published • 11 -
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
Paper • 2510.00406 • Published • 68 -
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
Paper • 2510.19430 • Published • 55
-
VibeVoice Technical Report
Paper • 2508.19205 • Published • 179 -
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
Paper • 2509.22186 • Published • 177 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 199 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 779
-
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
Paper • 2608.26103 • Published • 26 -
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
Paper • 2608.27550 • Published • 95 -
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Paper • 2608.27549 • Published • 53 -
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
Paper • 2608.25518 • Published • 196
-
WorldVLA: Towards Autoregressive Action World Model
Paper • 2506.21539 • Published • 40 -
LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
Paper • 2509.05263 • Published • 11 -
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
Paper • 2510.00406 • Published • 68 -
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
Paper • 2510.19430 • Published • 55
-
Wolf: Captioning Everything with a World Summarization Framework
Paper • 2407.18908 • Published • 33 -
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Paper • 2407.19985 • Published • 37 -
TPDiff: Temporal Pyramid Video Diffusion Model
Paper • 2503.09566 • Published • 45 -
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
Paper • 2506.07464 • Published • 14