VGPA · View-Graph Pose Averaging
Weizmann Institute of Science
Camera pose estimation is a key step in 3D reconstruction and view-synthesis pipelines. We present a deep, global Structure-from-Motion framework based on learned view-graph aggregation. Our method employs a permutation-equivariant, edge-conditioned graph neural network that takes noisy pairwise relative poses as input and outputs globally consistent camera extrinsics. The network is trained without ground-truth supervision, relying solely on a relative-pose consistency objective. This is followed by 3D point triangulation and robust bundle adjustment. Our approach is efficient, scalable to more than a thousand images, and robust to graph density. We evaluate our method on MegaDepth, 1DSfM, Strecha, and BlendedMVS. These experiments demonstrate that our method achieves superior rotation and translation accuracy compared to deep track-centric methods while registering more images across many scenes, and competitive results compared to state-of-the-art classical pipelines, while being much faster.
Runtime is the mean over the ten 1DSfM scenes, on the same point tracks as every baseline, and excludes the shared matching stage.
Interactive 3D views of the reconstructions. Press Explore in 3D on any of them to orbit it yourself.
The pipeline has two parts: preprocessing, which builds the view graph and the point tracks, and a network that averages the graph into global camera poses. The view graph has one node per image and one edge per relative pose recovered by RANSAC. Point tracks are used only afterwards, for triangulation and bundle adjustment.
SIFT matching and RANSAC give a partial set of essential matrices. Decomposing each yields a relative rotation Rij and a unit translation direction tij, which become the edges of the view graph. The matched pairs are also joined into point tracks, which are set aside until after the poses are solved.
All the geometry lies on the edges, so nodes start from a shared token and are updated by edge-conditioned message passing. The network is equivariant to relabelings of the graph.
A 3-layer MLP maps each node embedding to a translation and a quaternion, normalized to the unit sphere. One forward pass produces every camera in the scene at once.
The predicted cameras and the point tracks give 3D points by DLT triangulation, followed by robust bundle adjustment. An optional view-reintegration step adds back cameras that bundle adjustment discarded.
Nc is the number of input images, Nr the number registered. A marker highlights the best value in each row for each metric.
| Scene | Nc | Out.% | VGPA (ours) | RESfM | Theia | GLOMAP | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Nr | Rot° | Trans | Nr | Rot° | Trans | Nr | Rot° | Trans | Nr | Rot° | Trans | |||
| Alamo | 573 | 32.6 | 523 | 1.35 | 0.322 | 484 | 3.66 | 0.515 | 549 | 4.42 | 1.433 | 557 | 2.45 | 1.520 |
| Ellis Island | 227 | 25.1 | 215 | 0.28 | 0.081 | 214 | 0.82 | 0.122 | 213 | 5.01 | 1.527 | 219 | 0.58 | 0.155 |
| Madrid Metropolis | 333 | 39.4 | 298 | 1.31 | 0.145 | 244 | 8.42 | 0.827 | 319 | 2.61 | 0.903 | 320 | 1.22 | 0.242 |
| Montreal Notre Dame | 448 | 31.7 | 442 | 0.93 | 0.418 | 346 | 2.82 | 0.352 | 420 | 4.47 | 1.285 | 444 | 0.60 | 0.211 |
| NYC Library | 330 | 33.6 | 295 | 0.34 | 0.095 | 224 | 3.96 | 0.429 | 313 | 4.06 | 1.141 | 323 | 0.58 | 0.189 |
| Notre Dame | 549 | 35.6 | 527 | 0.52 | 0.105 | 517 | 1.20 | 0.231 | 531 | 3.70 | 0.828 | 543 | 2.73 | 0.389 |
| Piazza del Popolo | 336 | 33.1 | 318 | 5.37 | 0.670 | 249 | 2.20 | 0.186 | 324 | 3.31 | 1.053 | 331 | 0.80 | 0.188 |
| Tower of London | 467 | 27.0 | 457 | 0.89 | 0.057 | 94 | 0.67 | 0.026 | 448 | 6.61 | 1.189 | 466 | 0.81 | 0.138 |
| Vienna Cathedral | 824 | 31.4 | 763 | 7.90 | 0.590 | 479 | 1.52 | 0.112 | 767 | 12.25 | 1.663 | 822 | 2.00 | 2.414 |
| Yorkminster | 432 | 29.0 | 402 | 0.43 | 0.044 | 331 | 14.54 | 1.468 | 387 | 8.35 | 1.916 | 418 | 0.95 | 0.316 |
| Mean | 451 | 31.9 | 424 | 1.93 | 0.253 | 318 | 3.98 | 0.427 | 427 | 5.48 | 1.294 | 444 | 1.27 | 0.576 |
GLOMAP registers the most images on nine of ten scenes; VGPA is the most accurate on rotation and translation on the majority of them, and has the lowest mean translation error.
| Scene | Nc | Out.% | VGPA (ours) | RESfM | Theia | GLOMAP | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Nr | Rot° | Trans | Nr | Rot° | Trans | Nr | Rot° | Trans | Nr | Rot° | Trans | |||
| 0238 | 522 | 44.6 | 486 | 1.06 | 0.285 | 283 | 2.61 | 0.325 | 506 | 1.21 | 0.334 | 497 | 0.74 | 0.349 |
| 0060 | 528 | 41.6 | 518 | 0.10 | 0.026 | 503 | 0.29 | 0.029 | 525 | 0.85 | 0.124 | 520 | 0.11 | 0.048 |
| 0197 | 870 | 40.7 | 718 | 0.58 | 0.108 | 667 | 4.22 | 0.333 | 855 | 1.16 | 0.227 | 813 | 0.43 | 0.130 |
| 0094 | 763 | 40.1 | 659 | 0.81 | 0.116 | 537 | 3.77 | 0.750 | 742 | 0.75 | 0.160 | 711 | 0.88 | 3.907 |
| 0265 | 571 | 38.8 | 372 | 5.30 | 1.433 | 346 | 1.25 | 0.389 | 554 | 5.83 | 2.216 | 554 | 7.46 | 2.839 |
| 0083 | 635 | 31.3 | 622 | 0.32 | 0.062 | 596 | 0.64 | 0.058 | 632 | 0.37 | 0.372 | 614 | 0.08 | 0.016 |
| 0076 | 558 | 30.5 | 547 | 0.09 | 0.037 | 524 | 0.37 | 0.094 | 549 | 0.78 | 0.120 | 540 | 0.17 | 0.042 |
| 0185 | 368 | 30.0 | 364 | 0.12 | 0.021 | 350 | 0.06 | 0.010 | 365 | 0.41 | 0.094 | 365 | 0.16 | 0.051 |
| 0048 | 512 | 24.2 | 501 | 0.10 | 0.011 | 474 | 4.69 | 0.178 | 507 | 0.41 | 0.105 | 505 | 0.15 | 0.224 |
| 0024 | 356 | 23.0 | 328 | 3.59 | 1.122 | 309 | 2.03 | 0.398 | 355 | 0.56 | 0.219 | 338 | 0.15 | 0.104 |
| 0223 | 214 | 17.0 | 211 | 3.45 | 0.285 | 204 | 3.76 | 0.510 | 212 | 3.34 | 0.519 | 213 | 1.75 | 0.275 |
| 5016 | 28 | 16.9 | 28 | 0.08 | 0.015 | 28 | 0.12 | 0.016 | 28 | 0.10 | 0.061 | 28 | 0.08 | 0.046 |
| 0046 | 440 | 14.6 | 438 | 0.47 | 0.073 | 399 | 0.95 | 0.043 | 434 | 0.25 | 0.112 | 440 | 0.03 | 0.007 |
| Group 2 — Nc > 1000, subsampled to 300 | ||||||||||||||
| 0099 | 299 | 47.4 | 157 | 0.56 | 0.229 | 190 | 3.53 | 0.709 | 297 | 3.28 | 0.664 | 255 | 0.15 | 0.085 |
| 1001 | 285 | 43.9 | 268 | 3.87 | 2.585 | 251 | 1.70 | 0.661 | 275 | 7.89 | 3.990 | 270 | 4.56 | 3.817 |
| 0231 | 296 | 42.2 | 260 | 0.31 | 0.029 | 246 | 0.84 | 0.065 | 286 | 1.37 | 0.322 | 278 | 0.73 | 0.134 |
| 0411 | 299 | 29.9 | 268 | 0.23 | 0.036 | 273 | 0.13 | 0.020 | 293 | 0.39 | 0.196 | 268 | 0.19 | 0.148 |
| 0377 | 295 | 27.5 | 232 | 0.12 | 0.016 | 210 | 0.29 | 0.018 | 269 | 1.13 | 0.205 | 266 | 0.65 | 0.237 |
| 0102 | 299 | 25.8 | 296 | 0.18 | 0.023 | 284 | 0.28 | 0.059 | 294 | 2.31 | 0.698 | 293 | 0.15 | 0.101 |
| 0147 | 298 | 24.6 | 289 | 1.36 | 0.118 | 207 | 4.62 | 0.325 | 284 | 6.36 | 0.934 | 289 | 6.75 | 3.542 |
| 0148 | 287 | 24.6 | 244 | 1.15 | 0.135 | 197 | 0.60 | 0.035 | 275 | 13.98 | 1.558 | 282 | 22.73 | 2.646 |
| 0446 | 298 | 22.1 | 296 | 0.34 | 0.023 | 288 | 0.71 | 0.046 | 289 | 1.23 | 0.391 | 295 | 0.20 | 0.071 |
| 0022 | 297 | 21.2 | 280 | 0.11 | 0.015 | 274 | 0.29 | 0.039 | 296 | 0.58 | 0.160 | 280 | 0.22 | 0.087 |
| 0327 | 298 | 21.0 | 280 | 0.48 | 0.032 | 271 | 0.26 | 0.090 | 288 | 1.27 | 0.360 | 288 | 15.54 | 2.035 |
| 0015 | 284 | 20.6 | 241 | 0.61 | 0.083 | 215 | 1.04 | 0.167 | 244 | 2.21 | 0.389 | 272 | 0.28 | 0.095 |
| 0455 | 298 | 19.8 | 292 | 0.40 | 0.071 | 293 | 0.68 | 0.105 | 294 | 0.77 | 0.159 | 297 | 0.35 | 0.064 |
| 0496 | 297 | 19.2 | 280 | 0.69 | 0.041 | 281 | 0.35 | 0.055 | 285 | 1.40 | 0.550 | 291 | 0.44 | 0.303 |
| 1589 | 299 | 17.4 | 294 | 0.13 | 0.074 | 290 | 0.14 | 0.019 | 288 | 0.82 | 0.193 | 298 | 0.07 | 0.041 |
| 0012 | 299 | 16.3 | 295 | 0.77 | 0.086 | 287 | 0.40 | 0.027 | 129 | 1.04 | 0.318 | 295 | 0.51 | 0.121 |
| 0104 | 284 | 16.2 | 237 | 4.04 | 0.330 | 193 | 0.29 | 0.029 | 265 | 17.05 | 1.530 | 280 | 19.69 | 0.834 |
| 0019 | 299 | 15.4 | 288 | 0.23 | 0.008 | 250 | 0.06 | 0.008 | 271 | 0.81 | 0.250 | 296 | 0.09 | 0.025 |
| 0063 | 293 | 14.5 | 266 | 0.14 | 0.024 | 262 | 0.46 | 0.048 | 268 | 0.92 | 0.605 | 287 | 0.32 | 0.100 |
| 0130 | 285 | 14.4 | 202 | 0.21 | 0.019 | 192 | 0.20 | 0.023 | 187 | 1.20 | 0.349 | 279 | 2.00 | 0.909 |
| 0080 | 284 | 12.9 | 141 | 0.10 | 0.017 | 139 | 0.59 | 0.096 | 278 | 2.62 | 0.868 | 282 | 1.92 | 0.236 |
| 0240 | 298 | 11.9 | 288 | 0.64 | 0.088 | 275 | 3.13 | 0.265 | 285 | 1.31 | 0.470 | 290 | 0.39 | 0.135 |
| 0007 | 290 | 11.7 | 284 | 1.51 | 0.079 | 172 | 0.91 | 0.041 | 277 | 1.24 | 0.174 | 289 | 0.19 | 0.035 |
| Mean | 379 | 26.1 | 327 | 0.95 | 0.215 | 299 | 1.29 | 0.169 | 347 | 2.42 | 0.555 | 352 | 2.51 | 0.662 |
Above the divider are Group 1 scenes with fewer than 1000 images; below are Group 2 scenes with more than 1000 images, subsampled to 300 for testing.
| Scene | Nc | Out.% | VGPA (ours) | VGGT | MASt3R | VGGSfM | Theia | COLMAP | GLOMAP | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rot° | Trans | Time s | Rot° | Trans | Time s | Rot° | Trans | Time s | Rot° | Trans | Time s | Rot° | Trans | Time s | Rot° | Trans | Time s | Rot° | Trans | Time s | |||
| Strecha | |||||||||||||||||||||||
| entry-P10 | 10 | 4.8 | 0.004 | 0.0005 | 7 | 0.079 | 0.033 | 16.5 | 0.442 | 0.055 | 19 | 0.165 | 0.056 | 10.3 | 0.024 | 0.008 | 0.9 | 0.023 | 0.007 | 36.0 | 0.187 | 0.026 | 12.5 |
| fountain-P11 | 11 | 1.4 | 0.009 | 0.0003 | 8 | 0.034 | 0.019 | 12.2 | 0.160 | 0.026 | 22 | 0.172 | 0.016 | 15.4 | 0.027 | 0.002 | 1.5 | 0.027 | 0.003 | 37.0 | 0.194 | 0.022 | 38.6 |
| Herz-Jesus-P8 | 8 | 1.8 | 0.009 | 0.0010 | 6 | 0.032 | 0.011 | 12.7 | 0.363 | 0.037 | 16 | 0.206 | 0.042 | 8.7 | 0.025 | 0.005 | 0.6 | 0.026 | 0.004 | 22.0 | 0.091 | 0.015 | 5.0 |
| Herz-Jesus-P25 | 25 | 2.8 | 0.010 | 0.0003 | 12 | 0.048 | 0.007 | 31.9 | 0.869 | 0.057 | 81 | 0.158 | 0.046 | 19.6 | 0.026 | 0.006 | 2.4 | 0.028 | 0.006 | 60.0 | 0.138 | 0.013 | 76.6 |
| BlendedMVS | |||||||||||||||||||||||
| scene0 | 75 | 2.0 | 0.149 | 0.0204 | 73 | 0.041 | 0.017 | 108 | 0.501 | 0.191 | 516 | 0.045 | 0.0106 | 61 | 0.009 | 0.0017 | 49 | 0.006 | 0.0005 | 106 | 0.007 | 0.0016 | 198 |
| scene1 | 51 | 1.4 | 0.342 | 0.0345 | 28 | 0.101 | 0.050 | 41 | 0.919 | 0.173 | 1017 | 0.098 | 0.0112 | 32 | 0.029 | 0.0099 | 18 | 0.007 | 0.0003 | 67 | 0.024 | 0.0102 | 117 |
| scene2 | 33 | 2.2 | 0.008 | 0.0004 | 17 | 0.230 | 0.022 | 52 | 1.972 | 0.130 | 117 | 0.227 | 0.0180 | 30 | 0.045 | 0.0098 | 15 | 0.003 | 0.0002 | 55 | 0.025 | 0.0060 | 87 |
| scene3 | 66 | 8.8 | 0.006 | 0.0003 | 54 | 0.353 | 0.014 | 276 | 0.927 | 0.045 | 815 | 0.372 | 0.0174 | 52 | 0.019 | 0.0018 | 21 | 0.004 | 0.0002 | 128 | 0.008 | 0.0017 | 392 |
Small calibrated benchmarks with ground truth camera poses.
| Scene | Nc | VGPA (ours) | TTT3R | CUT3R | FAST3R | ||||
|---|---|---|---|---|---|---|---|---|---|
| Rot° | Trans | Rot° | Trans | Rot° | Trans | Rot° | Trans | ||
| Alamo | 573 | 1.35 | 0.322 | 16.08 | 3.650 | 22.32 | 4.266 | 40.37 | 3.990 |
| Ellis Island | 227 | 0.28 | 0.081 | 8.85 | 2.333 | 12.53 | 3.171 | 11.74 | 3.201 |
| Madrid Metropolis | 333 | 1.31 | 0.145 | 14.83 | 2.258 | 18.81 | 3.189 | 67.37 | 3.574 |
| Montreal Notre Dame | 448 | 0.93 | 0.418 | 13.25 | 1.230 | 16.12 | 2.134 | 10.79 | 2.052 |
| NYC Library | 330 | 0.34 | 0.095 | 7.67 | 1.656 | 10.26 | 1.912 | 11.90 | 2.408 |
| Notre Dame | 549 | 0.52 | 0.105 | 11.88 | 1.430 | 15.33 | 1.694 | 16.44 | 2.241 |
| Piazza del Popolo | 336 | 5.37 | 0.670 | 23.13 | 2.063 | 22.63 | 2.092 | 32.48 | 2.568 |
| Tower of London | 467 | 0.89 | 0.057 | 29.69 | 3.647 | 29.85 | 3.607 | 59.53 | 3.658 |
| Vienna Cathedral | 824 | 7.90 | 0.590 | 43.76 | 2.978 | 41.58 | 2.802 | 29.49 | 2.561 |
| Yorkminster | 432 | 0.43 | 0.044 | 16.45 | 2.223 | 23.00 | 2.992 | 20.26 | 2.463 |
Feed-forward 3D reconstruction models on the 1DSfM scenes. Errors for these models are an order of magnitude larger, and they perform badly on this type of data.
| Scene | Nc | Exhaustive matching | MegaLoc top-30 | ||||
|---|---|---|---|---|---|---|---|
| Nr | Rot° | Trans | Nr | Rot° | Trans | ||
| Alamo | 573 | 523 | 1.35 | 0.322 | 548 | 1.18 | 0.163 |
| Ellis Island | 227 | 215 | 0.28 | 0.081 | 216 | 0.21 | 0.090 |
| Madrid Metropolis | 333 | 298 | 1.31 | 0.145 | 323 | 1.48 | 3.397 |
| Montreal Notre Dame | 448 | 442 | 0.93 | 0.418 | 338 | 0.11 | 0.015 |
| NYC Library | 330 | 295 | 0.34 | 0.095 | 319 | 0.92 | 0.537 |
| Notre Dame | 549 | 527 | 0.52 | 0.105 | 535 | 0.42 | 0.086 |
| Piazza del Popolo | 336 | 318 | 5.37 | 0.670 | 329 | 0.57 | 0.064 |
| Tower of London | 467 | 457 | 0.89 | 0.057 | 447 | 2.14 | 0.408 |
| Vienna Cathedral | 824 | 763 | 7.90 | 0.590 | 771 | 9.32 | 0.462 |
| Yorkminster | 432 | 402 | 0.43 | 0.044 | 405 | 6.32 | 0.469 |
Replacing exhaustive pairwise matching with a sparse top-30 MegaLoc retrieval graph. Despite the large drop in edge density, VGPA holds registration coverage and accuracy.
| Scene | Nc | Out.% | Ours — intrinsics known | Ours — intrinsics estimated | ||||
|---|---|---|---|---|---|---|---|---|
| Nr | Rot° | Trans | Nr | Rot° | Trans | |||
| BlendedMVS — shared intrinsics | ||||||||
| scene0 | 75 | 2.0 | 74 | 0.149 | 0.0204 | 74 | 0.122 | 0.0159 |
| scene1 | 51 | 1.4 | 51 | 0.342 | 0.0345 | 51 | 0.337 | 0.0401 |
| scene2 | 33 | 2.2 | 33 | 0.008 | 0.0004 | 33 | 0.025 | 0.0050 |
| scene3 | 66 | 8.8 | 66 | 0.006 | 0.0003 | 66 | 0.010 | 0.0014 |
| MegaDepth — intrinsics not shared | ||||||||
| 0012 | 299 | 16.3 | 295 | 0.77 | 0.086 | 292 | 0.98 | 0.253 |
| 0024 | 356 | 23.0 | 328 | 3.59 | 1.122 | 308 | 1.79 | 0.508 |
| 0048 | 512 | 24.2 | 501 | 0.10 | 0.011 | 485 | 0.38 | 0.456 |
| 0083 | 635 | 31.3 | 622 | 0.32 | 0.062 | 596 | 0.64 | 0.058 |
Calibration is not required.
Reconstructions and recovered cameras across the different datasets. Click any thumbnail to enlarge.
@inproceedings{khatib2026vgpa,
title = {Learning Global Camera Poses from Noisy View-Graphs
for Structure from Motion},
author = {Khatib, Fadi and Galun, Meirav and Basri, Ronen},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026}
}