-
sensenova/SenseNova-U1-8B-MoT
Any-to-Any • 18B • Updated • 22.8k • 291 -
sensenova/SenseNova-U1-8B-MoT-Infographic-V3
Any-to-Any • 18B • Updated • 8.63k • 60 -
sensenova/SenseNova-U1-8B-MoT-Infographic-V2
Any-to-Any • 18B • Updated • 50 • 29 -
sensenova/SenseNova-U1-8B-MoT-Infographic
Any-to-Any • 18B • Updated • 55 • 56
Collections
Discover the best community collections!
Collections including paper arxiv:2605.12500
-
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
Paper • 2607.28618 • Published • 304 -
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Paper • 2607.28227 • Published • 311 -
Metis: Memory Foundation Model
Paper • 2607.26760 • Published • 274 -
Kimi K3: Open Frontier Intelligence
Paper • 2607.24653 • Published • 514
-
Code as Agent Harness
Paper • 2605.18747 • Published • 227 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 199 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 172 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 145
-
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Paper • 2603.25746 • Published • 156 -
TAPS: Task Aware Proposal Distributions for Speculative Sampling
Paper • 2603.27027 • Published • 148 -
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Paper • 2603.25716 • Published • 158 -
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Paper • 2603.27538 • Published • 151
-
VibeVoice Technical Report
Paper • 2508.19205 • Published • 177 -
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
Paper • 2509.22186 • Published • 177 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 199 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 772
-
GLM-5: from Vibe Coding to Agentic Engineering
Paper • 2602.15763 • Published • 219 -
zai-org/GLM-5.2
Text Generation • 753B • Updated • 997k • • 5.07k -
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Paper • 2605.23904 • Published • 265 -
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
Paper • 2503.11576 • Published • 174
-
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 199 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 172 -
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
Paper • 2605.03849 • Published • 130 -
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Paper • 2605.03042 • Published • 150
-
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Paper • 2412.03069 • Published • 34 -
Are Emergent Abilities of Large Language Models a Mirage?
Paper • 2304.15004 • Published • 8 -
Scaling Image Tokenizers with Grouped Spherical Quantization
Paper • 2412.02632 • Published • 10 -
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Paper • 2410.13848 • Published • 37
-
sensenova/SenseNova-U1-8B-MoT
Any-to-Any • 18B • Updated • 22.8k • 291 -
sensenova/SenseNova-U1-8B-MoT-Infographic-V3
Any-to-Any • 18B • Updated • 8.63k • 60 -
sensenova/SenseNova-U1-8B-MoT-Infographic-V2
Any-to-Any • 18B • Updated • 50 • 29 -
sensenova/SenseNova-U1-8B-MoT-Infographic
Any-to-Any • 18B • Updated • 55 • 56
-
VibeVoice Technical Report
Paper • 2508.19205 • Published • 177 -
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
Paper • 2509.22186 • Published • 177 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 199 -
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
Paper • 2608.09888 • Published • 772
-
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
Paper • 2607.28618 • Published • 304 -
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Paper • 2607.28227 • Published • 311 -
Metis: Memory Foundation Model
Paper • 2607.26760 • Published • 274 -
Kimi K3: Open Frontier Intelligence
Paper • 2607.24653 • Published • 514
-
GLM-5: from Vibe Coding to Agentic Engineering
Paper • 2602.15763 • Published • 219 -
zai-org/GLM-5.2
Text Generation • 753B • Updated • 997k • • 5.07k -
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Paper • 2605.23904 • Published • 265 -
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
Paper • 2503.11576 • Published • 174
-
Code as Agent Harness
Paper • 2605.18747 • Published • 227 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 199 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 172 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 145
-
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 199 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 172 -
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
Paper • 2605.03849 • Published • 130 -
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Paper • 2605.03042 • Published • 150
-
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Paper • 2603.25746 • Published • 156 -
TAPS: Task Aware Proposal Distributions for Speculative Sampling
Paper • 2603.27027 • Published • 148 -
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Paper • 2603.25716 • Published • 158 -
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Paper • 2603.27538 • Published • 151
-
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Paper • 2412.03069 • Published • 34 -
Are Emergent Abilities of Large Language Models a Mirage?
Paper • 2304.15004 • Published • 8 -
Scaling Image Tokenizers with Grouped Spherical Quantization
Paper • 2412.02632 • Published • 10 -
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Paper • 2410.13848 • Published • 37