Accurate Surgical Scene Reconstruction from Multi-view FoundationStereo
DOI:
https://doi.org/10.54097/y6fzh146Keywords:
Computer Vision, 3D Reconstruction, Truncated Signed Distance Function (TSDF), Foundation Stereo, Multiview Stereo.Abstract
Accurate 3D reconstruction of surgical scenes is a critical enabling technology for advancements in intraoperative navigation, surgical training, and robotic automation. While learning-based stereo depth estimation methods have demonstrated high in-domain accuracy, their performance often degrades significantly under domain shifts. This limitation is particularly acute in surgical applications, where large-scale, annotated datasets are scarce. In this work, we investigate the application of FoundationStereo, a recently proposed vision foundation model for stereo matching, to the task of surgical scene reconstruction. We leverage its zero-shot, single-frame depth estimation capabilities within a multi-view fusion framework based on the Truncated Signed Distance Function (TSDF) to achieve comprehensive scene reconstruction. Our experiments, conducted on the public SCARED dataset captured with a da Vinci Xi surgical robot, demonstrate that FoundationStereo achieves state-of-the-art zero-shot accuracy. We report a sub-millimeter mean error for single-frame depth estimation and an error under 2 mm for fused multi-view reconstructions, significantly outperforming the zero-shot STTR baseline. These results highlight the substantial potential of FoundationStereo for enabling accurate, high-fidelity surgical scene reconstruction without domain-specific training. We also discuss current limitations, including the reliance on known camera pose information and challenges in dynamic scenes, and outline future research directions to enhance robustness for clinical applications.
References
[1] F. Liu, C. Shen, G. Lin, and I. Reid, “Learning depth from single monocular images using deep convolutional neural fields,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 38, no. 10, pp. 2024–2039, Oct. 2016, doi: 10.1109/tpami.2015.2505283.
[2] Z. Li et al., “Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 6197–6206.
[3] B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis.” 2020. Available: https://arxiv.org/abs/2003.08934
[4] B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3D gaussian splatting for real-time radiance field rendering.” 2023. Available: https://arxiv.org/abs/2308.04079
[5] B. D. Killeen et al., “Stand in surgeon’s shoes: Virtual reality cross-training to enhance teamwork in surgery,” International journal of computer assisted radiology and surgery, vol. 19, no. 6, pp. 1213–1222, 2024.
[6] H. Ding et al., “Towards robust automation of surgical systems via digital twin-based scene representations from foundation models,” arXiv preprint arXiv:2409.13107, 2024.
[7] C. Godard, O. Mac Aodha, M. Firman, and G. J. Brostow, “Digging into self-supervised monocular depth estimation,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 3828–3838.
[8] A. Kirillov et al., “Segment anything,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4015–4026.
[9] N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht, “Cotracker: It is better to track together,” in European conference on computer vision, Springer, 2024, pp. 18–35.
[10] A. Bochkovskii et al., “Depth pro: Sharp monocular metric depth in less than a second.” 2025. Available: https://arxiv.org/abs/2410.02073
[11] L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data.” 2024. Available: https://arxiv.org/abs/2401.10891
[12] B. Wen, M. Trepte, J. Aribido, J. Kautz, O. Gallo, and S. Birchfield, “FoundationStereo: Zero-shot stereo matching.” 2025. Available: https://arxiv.org/abs/2501.09898
[13] M. Grinvald, F. Tombari, R. Siegwart, and J. Nieto, “TSDF++: A multi-object formulation for dynamic object tracking and reconstruction.” 2021. Available: https://arxiv.org/abs/2105.07468
[14] M. Allan et al., “Stereo correspondence and reconstruction of endoscopic data challenge.” 2021. Available: https://arxiv.org/abs/2101.01133
[15] D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network.” 2014. Available: https://arxiv.org/abs/1406.2283
[16] I. Laina, C. Rupprecht, V. Belagiannis, F. Tombari, and N. Navab, “Deeper depth prediction with fully convolutional residual networks.” 2016. Available: https://arxiv.org/abs/1606.00373
[17] R. Ranftl, A. Bochkovskiy, and V. Koltun, “Vision transformers for dense prediction.” 2021. Available: https://arxiv.org/abs/2103.13413
[18] A. Saxena, S. Chung, and A. Ng, “Learning depth from single monocular images,” in Advances in neural information processing systems, Y. Weiss, B. Schölkopf, and J. Platt, Eds., MIT Press, 2005. Available: https://proceedings.neurips.cc/paper_files/paper/2005/file/17d8da815fa21c57af9829fb0a869602-Paper.pdf
[19] A. Saxena, M. Sun, and A. Y. Ng, “Make3D: Learning 3D scene structure from a single still image,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 31, no. 5, pp. 824–840, 2009, doi: 10.1109/TPAMI.2008.132.
[20] S. F. Bhat, I. Alhashim, and P. Wonka, “LocalBins: Improving depth estimation by learning local distributions.” 2022. Available: https://arxiv.org/abs/2203.15132
[21] S. Farooq Bhat, I. Alhashim, and P. Wonka, “AdaBins: Depth estimation using adaptive bins,” in 2021 IEEE/CVF conference on computer vision and pattern recognition (CVPR), IEEE, Jun. 2021, pp. 4008–4017. doi: 10.1109/cvpr46437.2021.00400.
[22] Y. Wang, Z. Pan, X. Li, Z. Cao, K. Xian, and J. Zhang, “Less is more: Consistent video depth estimation with masked frames modeling,” in Proceedings of the 30th ACM international conference on multimedia, in MM ’22. New York, NY, USA: Association for Computing Machinery, 2022, pp. 6347–6358. doi: 10.1145/3503161.3547978.
[23] J. Choe, S. Im, F. Rameau, M. Kang, and I. S. Kweon, “VolumeFusion: Deep depth fusion for 3D scene reconstruction.” 2021. Available: https://arxiv.org/abs/2108.08623
[24] Y. Furukawa and C. Hernández, “Multi-view stereo: A tutorial,” Foundations and Trends® in Computer Graphics and Vision, vol. 9, no. 1–2, pp. 1–148, 2015, doi: 10.1561/0600000052.
[25] Q. Fu, Q. Xu, Y.-S. Ong, and W. Tao, “Geo-neus: Geometry-consistent neural implicit surfaces learning for multi-view reconstruction.” 2022. Available: https://arxiv.org/abs/2205.15848
[26] M. Oechsle, S. Peng, and A. Geiger, “UNISURF: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction.” 2021. Available: https://arxiv.org/abs/2104.10078
[27] M. Oquab et al., “DINOv2: Learning robust visual features without supervision.” 2024. Available: https://arxiv.org/abs/2304.07193
[28] Y. Wang, I. Skorokhodov, and P. Wonka, “HF-NeuS: Improved surface reconstruction using high-frequency details.” 2022. Available: https://arxiv.org/abs/2206.07850
[29] P. Wang, L. Liu, Y. Liu, C. Theobalt, T. Komura, and W. Wang, “NeuS: Learning neural implicit surfaces by volume rendering for multi-view reconstruction.” 2023. Available: https://arxiv.org/abs/2106.10689
[30] Y. Yao, Z. Luo, S. Li, T. Fang, and L. Quan, “MVSNet: Depth inference for unstructured multi-view stereo.” 2018. Available: https://arxiv.org/abs/1804.02505
[31] X. Ye, W. Zhao, T. Liu, Z. Huang, Z. Cao, and X. Li, “Constraining depth map geometry for multi-view stereo: A dual-depth approach with saddle-shaped depth cells.” 2023. Available: https://arxiv.org/abs/2307.09160
[32] C. Zhao, Y. Ge, F. Zhu, R. Zhao, H. Li, and M. Salzmann, “Progressive correspondence pruning by consensus learning.” 2021. Available: https://arxiv.org/abs/2101.00591
[33] S. Wang, V. Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud, “DUSt3R: Geometric 3D vision made easy.” 2024. Available: https://arxiv.org/abs/2312.14132
[34] Z. Li and N. Snavely, “MegaDepth: Learning single-view depth prediction from internet photos,” in 2018 IEEE/CVF conference on computer vision and pattern recognition, 2018, pp. 2041–2050. doi: 10.1109/CVPR.2018.00218.
[35] R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,” IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2020.
[36] S. Li, Y. Luo, Y. Zhu, X. Zhao, Y. Li, and Y. Shan, “Enforcing temporal consistency in video depth estimation,” in Proceedings of the IEEE/CVF international conference on computer vision (ICCV) workshops, 2021, pp. 1145–1154.
[37] H. Zhang, C. Shen, Y. Li, Y. Cao, Y. Liu, and Y. Yan, “Exploiting temporal consistency for real-time video depth estimation.” 2019. Available: https://arxiv.org/abs/1908.03706
[38] R. Birkl, D. Wofk, and M. Müller, “MiDaS v3.1 – a model zoo for robust monocular relative depth estimation.” 2023. Available: https://arxiv.org/abs/2307.14460
[39] A. Petrovai and S. Nedevschi, “Exploiting pseudo labels in a self-supervised learning framework for improved monocular depth estimation,” in 2022 IEEE/CVF conference on computer vision and pattern recognition (CVPR), 2022, pp. 1568–1578. doi: 10.1109/CVPR52688.2022.00163.
[40] J. Spencer, C. Russell, S. Hadfield, and R. Bowden, “Kick back & relax: Learning to reconstruct the world by watching SlowTV.” 2023. Available: https://arxiv.org/abs/2307.10713
[41] M. Gui et al., “DepthFM: Fast monocular depth estimation with flow matching.” 2024. Available: https://arxiv.org/abs/2403.13788
[42] B. Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler, “Repurposing diffusion-based image generators for monocular depth estimation.” 2024. Available: https://arxiv.org/abs/2312.02145
[43] H. Fu, M. Gong, C. Wang, K. Batmanghelich, and D. Tao, “Deep ordinal regression network for monocular depth estimation.” 2018. Available: https://arxiv.org/abs/1806.02446
[44] S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. Müller, “ZoeDepth: Zero-shot transfer by combining relative and metric depth.” 2023. Available: https://arxiv.org/abs/2302.12288
[45] R. Zhu, C. Wang, Z. Song, L. Liu, T. Zhang, and Y. Zhang, “ScaleDepth: Decomposing metric depth estimation into scale prediction and relative depth estimation.” 2024. Available: https://arxiv.org/abs/2407.08187
[46] M. Hu et al., “Metric3D v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 10579–10596, Dec. 2024, doi: 10.1109/tpami.2024.3444912.
[47] W. Yin et al., “Metric3D: Towards zero-shot metric 3D prediction from a single image.” 2023. Available: https://arxiv.org/abs/2307.10984
[48] L. Piccinelli et al., “UniDepth: Universal monocular metric depth estimation.” 2024. Available: https://arxiv.org/abs/2403.18913
[49] L. Yang et al., “Depth anything V2.” 2024. Available: https://arxiv.org/abs/2406.09414
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







