NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis
Abstract
Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3 times faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis. Additional qualitative results, videos, and resources are available at https://corl-team.github.io/namvis/
Community
NAMVIS replaces diffusion with next-scale autoregression for sparse-view multi-view synthesis: target views are generated coarse-to-fine in 7 scale steps, with all tokens of a scale predicted in parallel and a multi-scale PRoPE injecting camera geometry into attention. It takes 0.6 s per view vs 2.0 s for the fastest diffusion baseline we tested, with better PSNR/SSIM/LPIPS on Objaverse, GSO and OmniObject3D. Code, the 1B checkpoint and ~200K Objaverse-XL renders are open.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis (2026)
- GenRec: Knowing Where to Reconstruct and Where to Generate (2026)
- Sparse auto-regressive modeling for scene generation from multi-view images (2026)
- RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting (2026)
- UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing (2026)
- DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion (2026)
- GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.04722 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash