Scaffold3D

SfM-Conditioned Pointmap Prediction
for Multi-View 3D Reconstruction

1ETH Zurich    2Microsoft

TL;DR: Scaffold3D turns an SfM “scaffold” (cameras + 3D points) into patch-aligned point tokens and injects them into image-token decoding for more accurate multi-view 3D reconstruction.

ETH3D Cross-domain Facade scene with a large SfM-conditioning gain.
Conditioning-free (VGGT)
SfM scaffold
Scaffold3D

Abstract

Multi-view 3D reconstruction is increasingly driven by feed-forward models, which are fast and robust but often imprecise or globally inconsistent across views, especially in sparse-view and low-overlap settings. Structure-from-motion (SfM) offers a complementary geometric signal through camera poses and 3D points estimated by correspondence filtering and optimization. To propagate this globally consistent SfM scaffold into dense geometry, we introduce Scaffold3D, an SfM-conditioned reconstruction framework that injects image patch-aligned SfM point tokens into a pairwise feed-forward pointmap predictor. Our approach preserves image-token reasoning while using explicit SfM structure to condition dense prediction. Since pairwise pointmaps are predicted in local frames, we fuse them with global alignment and an SfM anchoring term. Across established benchmarks (ScanNet++, ETH3D, and Tanks&Temples) and out-of-domain data (4D-DRESS and MV-dVRK), Scaffold3D achieves stronger overall performance than feed-forward reconstruction models, point-token propagation, depth-completion baselines, and recent geometry-conditioned models. The gains are especially clear in low- and no-overlap evaluations. Notably, our two-view SfM-conditioned pointmap predictor outperforms several off-the-shelf multi-view geometry-conditioned baselines. Together, these findings establish SfM-scaffold conditioning as a practical interface between learned pointmap prediction and classical geometric optimization.

Key insights

1. Start from an SfM scaffold Feed-forward models are dense but can be inconsistent; an SfM scaffold, i.e., cameras and 3D points, is reliable where matching and triangulation succeed.
2. Keep image-token decoding Unlike point-only propagation, Scaffold3D conditions dense pointmap prediction with SfM point tokens while preserving visual cues in the image-token stream.
3. Anchor to scaffold SfM-anchored global alignment after the dense prediction fuses pairwise pointmaps into a coherent scene without drifting away from the optimized scaffold.

How Scaffold3D differs from related methods

Feed-forward predictors DUSt3R, MASt3R-SfM, VGGT, and Pi3 predict dense geometry from images first, then optionally align or refine it afterward. Scaffold3D instead feeds SfM cameras and points into the dense predictor while it estimates the pointmaps. Tables 1, 2, 3.
Monocular depth completion OMNI-DC and Marigold-DC complete each view independently from scaffold depth samples. Scaffold3D uses the same scaffold as multi-view conditioning, so other views can help where direct triangulation is missing. Tables 1, A4, A5.
Point-token propagation GGPT turns the scaffold into point tokens and refines a 3D point cloud. Scaffold3D keeps explicit point tokens too, but feeds them into the image-token decoder, so RGB appearance and matching cues remain active. Tables 1, 2, 3, A4, A5.
Multi-view stereo MVSAnywhere uses scaffold cameras and point-derived depth ranges to estimate view depths. It gets camera/depth-range input, but not the triangulated scaffold points as tokens inside the predictor. Tables 1, 2, A4, A5.
Geometry-conditioned predictors Pow3R, MapAnything, and Pi3X already condition dense predictors with SfM-derived depth, camera, or pose features. Scaffold3D studies a different interface: it pools scaffold observations into patch-aligned 3D point tokens, refines those tokens with self-attention, and injects them into image-token decoding. Tables 1, 2, A2, A4, A5.

Scaffold3D in 5 minutes

Method

Scaffold3D is a two-view SfM-conditioned pointmap predictor with a lightweight point encoder. For each image patch, all SfM 3D points whose projections fall into that patch are encoded together with their within-patch offsets, aggregated into a patch-aligned point token, refined by self-attention, and added to the image-token stream before decoding. This preserves dense visual reasoning while letting the SfM scaffold guide predictions beyond directly triangulated support.

Scaffold3D method overview
STEP 1: Build the SfM Scaffold

We first build an SfM scaffold from the input views. Dense correspondences are matched and filtered, bundle adjustment estimates cameras and feature tracks, and the surviving correspondences are triangulated into image-aligned 3D points.

1.1 Input Views
Input RGB view 1 Input RGB view 2 Input RGB view 3 Input RGB view 4
1.2 Matching (RoMa V2)
views 1 and 2
Filtered RoMa rainbow matches in view 1 for pair 1 and 2 Filtered RoMa rainbow matches in view 2 for pair 1 and 2
views 1 and 3
Filtered RoMa rainbow matches in view 1 for pair 1 and 3 Filtered RoMa rainbow matches in view 3 for pair 1 and 3
views 1 and 4
Filtered RoMa rainbow matches in view 1 for pair 1 and 4 Filtered RoMa rainbow matches in view 4 for pair 1 and 4
views 2 and 3
Filtered RoMa rainbow matches in view 2 for pair 2 and 3 Filtered RoMa rainbow matches in view 3 for pair 2 and 3
views 2 and 4
Filtered RoMa rainbow matches in view 2 for pair 2 and 4 Filtered RoMa rainbow matches in view 4 for pair 2 and 4
views 3 and 4
Filtered RoMa rainbow matches in view 3 for pair 3 and 4 Filtered RoMa rainbow matches in view 4 for pair 3 and 4
1.3 Bundle Adjustment
1.4 Direct Linear Triangulation
Final triangulated scaffold overlay
STEP 2: SfM-Conditioned Pointmap Prediction

For each view pair, Scaffold3D predicts complete pointmaps from RGB images and their SfM scaffolds. SfM observations in the same image patch are encoded as a patch-aligned point token and added to the image-token decoder, so dense visual reasoning stays active while the scaffold guides geometry beyond directly triangulated regions.

Views 1 | 2
STEP 3: SfM-Anchored Global Alignment
Open in Rerun

Fixed-camera dense global alignment fuses the pairwise pointmaps into one multi-view reconstruction. Pairwise terms complete dense geometry, and SfM anchors constrain pixels with triangulated scaffold support.

3D error (no scaffold)
SfM scaffold (fixed)
3D error (full scene)
3D error (close-up)
0% Loading Step 3 videos before synced playback...

Qualitative comparison

The paper evaluates Scaffold3D against feed-forward predictors (DUSt3R, MASt3R, MASt3R-SfM, VGGT, Pi3, Pi3X, Pow3R, and MapAnything), monocular depth-completion methods (OMNI-DC and Marigold-DC), point-token propagation (GGPT), and geometry-conditioned predictors (MVSAnywhere, MapAnything, Pi3X, and Pow3R) under a shared SfM scaffold. In the interactive viewer below, any evaluation scene from any dataset can be selected to compare full point clouds from a compact set of representative methods. All conditioning-based methods are given the same scaffold for each scene.

ETH3D Cross-domain Outdoor geometry with large planar structures.

Limitations

Scaffold3D assumes the SfM scaffold is reliable where it exists. This held on all five main benchmarks: the epipolar, reprojection, and cycle-consistency filters reject gross errors, so difficult scenes mostly lose coverage rather than gain wrong points. The exception is textureless surfaces under extreme perspective change, as in out-of-domain hand-object scenes, where matches drift along the epipolar line, pass every filter, and triangulate to the wrong depth. Conditioning still improves our backbone there, but a feed-forward model with a stronger domain prior starts higher.

DexYCB Limitation 4B hand-object scene with a different subject and object layout.

Acknowledgements

This study was conducted within the national “Proficiency” research project (No. PFFS-21-19), funded by the Swiss Innovation Agency Innosuisse in 2021 as one of 15 flagship initiatives. This work was supported by the Swiss AI Initiative, with computational resources provided by the Swiss National Supercomputing Centre (CSCS) on the Alps infrastructure. We thank Luigi Piccinelli, Yiming Wang, Sergey Prokudin, Nikita Karaev, and Xudong Jiang for valuable discussions and support.

BibTeX

@inproceedings{rajic2026scaffold3d,
  title     = {Scaffold3D: SfM-Conditioned Pointmap Prediction for Multi-View 3D Reconstruction},
  author    = {Raji{\v{c}}, Frano and Chen, Yutong and Xu, Haofei and Pataki, Zador and Pollefeys, Marc and Tang, Siyu},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}