Multi-view 3D reconstruction is increasingly driven by feed-forward models, which are fast and robust but often imprecise or globally inconsistent across views, especially in sparse-view and low-overlap settings. Structure-from-motion (SfM) offers a complementary geometric signal through camera poses and 3D points estimated by correspondence filtering and optimization. To propagate this globally consistent SfM scaffold into dense geometry, we introduce Scaffold3D, an SfM-conditioned reconstruction framework that injects image patch-aligned SfM point tokens into a pairwise feed-forward pointmap predictor. Our approach preserves image-token reasoning while using explicit SfM structure to condition dense prediction. Since pairwise pointmaps are predicted in local frames, we fuse them with global alignment and an SfM anchoring term. Across established benchmarks (ScanNet++, ETH3D, and Tanks&Temples) and out-of-domain data (4D-DRESS and MV-dVRK), Scaffold3D achieves stronger overall performance than feed-forward reconstruction models, point-token propagation, depth-completion baselines, and recent geometry-conditioned models. The gains are especially clear in low- and no-overlap evaluations. Notably, our two-view SfM-conditioned pointmap predictor outperforms several off-the-shelf multi-view geometry-conditioned baselines. Together, these findings establish SfM-scaffold conditioning as a practical interface between learned pointmap prediction and classical geometric optimization.
Scaffold3D is a two-view SfM-conditioned pointmap predictor with a lightweight point encoder. For each image patch, all SfM 3D points whose projections fall into that patch are encoded together with their within-patch offsets, aggregated into a patch-aligned point token, refined by self-attention, and added to the image-token stream before decoding. This preserves dense visual reasoning while letting the SfM scaffold guide predictions beyond directly triangulated support.
We first build an SfM scaffold from the input views. Dense correspondences are matched and filtered, bundle adjustment estimates cameras and feature tracks, and the surviving correspondences are triangulated into image-aligned 3D points.
For each view pair, Scaffold3D predicts complete pointmaps from RGB images and their SfM scaffolds. SfM observations in the same image patch are encoded as a patch-aligned point token and added to the image-token decoder, so dense visual reasoning stays active while the scaffold guides geometry beyond directly triangulated regions.
Fixed-camera dense global alignment fuses the pairwise pointmaps into one multi-view reconstruction. Pairwise terms complete dense geometry, and SfM anchors constrain pixels with triangulated scaffold support.
The paper evaluates Scaffold3D against feed-forward predictors (DUSt3R, MASt3R, MASt3R-SfM, VGGT, Pi3, Pi3X, Pow3R, and MapAnything), monocular depth-completion methods (OMNI-DC and Marigold-DC), point-token propagation (GGPT), and geometry-conditioned predictors (MVSAnywhere, MapAnything, Pi3X, and Pow3R) under a shared SfM scaffold. In the interactive viewer below, any evaluation scene from any dataset can be selected to compare full point clouds from a compact set of representative methods. All conditioning-based methods are given the same scaffold for each scene.
Scaffold3D assumes the SfM scaffold is reliable where it exists. This held on all five main benchmarks: the epipolar, reprojection, and cycle-consistency filters reject gross errors, so difficult scenes mostly lose coverage rather than gain wrong points. The exception is textureless surfaces under extreme perspective change, as in out-of-domain hand-object scenes, where matches drift along the epipolar line, pass every filter, and triangulate to the wrong depth. Conditioning still improves our backbone there, but a feed-forward model with a stronger domain prior starts higher.
This study was conducted within the national “Proficiency” research project (No. PFFS-21-19), funded by the Swiss Innovation Agency Innosuisse in 2021 as one of 15 flagship initiatives. This work was supported by the Swiss AI Initiative, with computational resources provided by the Swiss National Supercomputing Centre (CSCS) on the Alps infrastructure. We thank Luigi Piccinelli, Yiming Wang, Sergey Prokudin, Nikita Karaev, and Xudong Jiang for valuable discussions and support.
@inproceedings{rajic2026scaffold3d,
title = {Scaffold3D: SfM-Conditioned Pointmap Prediction for Multi-View 3D Reconstruction},
author = {Raji{\v{c}}, Frano and Chen, Yutong and Xu, Haofei and Pataki, Zador and Pollefeys, Marc and Tang, Siyu},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026}
}