BlendedMVS: A Large-scale Dataset for Generalized Multi-view Stereo Networks

被引：294

作者：

Yao, Yao ^{[1
]}

Luo, Zixin ^{[1
]}

Li, Shiwei ^{[2
]}

Zhang, Jingyang ^{[1
]}

Ren, Yufan ^{[3
]}

Zhou, Lei ^{[1
]}

Fang, Tian ^{[2
]}

Quan, Long ^{[1
]}

机构：

[1] Hong Kong Univ Sci & Technol, Hong Kong, Peoples R China

[2] Everest Innovat Technol, Hong Kong, Peoples R China

[3] Zhejiang Univ, Hangzhou, Zhejiang, Peoples R China

来源：

2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR) | 2020年

关键词：

D O I：

10.1109/CVPR42600.2020.00186

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

While deep learning has recently achieved great success on multi-view stereo (MVS), limited training data makes the trained model hard to be generalized to unseen scenarios. Compared with other computer vision tasks, it is rather difficult to collect a large-scale MVS dataset as it requires expensive active scanners and labor-intensive process to obtain ground truth 3D structures. In this paper, we introduce BlendedMVS, a novel large-scale dataset, to provide sufficient training ground truth for learning-based MVS. To create the dataset, we apply a 3D reconstruction pipeline to recover high-quality textured meshes from images of well-selected scenes. Then, we render these mesh models to color images and depth maps. To introduce the ambient lighting information during training, the rendered color images are further blended with the input images to generate the training input. Our dataset contains over 17k high-resolution images covering a variety of scenes, including cities, architectures, sculptures and small objects. Extensive experiments demonstrate that BlendedMVS endows the trained model with significantly better generalization ability compared with other MVS datasets. The dataset and pretrained models are available at https://github.com/YoYo000/BlendedMVS.

引用

页码：1787 / 1796

页数：10

共 31 条

[11] SurfaceNet: An End-to-end 3D Neural Network for Multiview Stereopsis [J].

Ji, Mengqi ;

Gall, Juergen ;

Zheng, Haitian ;

Liu, Yebin ;

Fang, Lu .

2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, :2326-2334

[12] End-to-End Learning of Geometry and Context for Deep Stereo Regression [J].

Kendall, Alex ;

Martirosyan, Hayk ;

Dasgupta, Saumitro ;

Henry, Peter ;

Kennedy, Ryan ;

Bachrach, Abraham ;

Bry, Adam .

2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, :66-75

[13]

Keyang Luo, 2019, INT C COMP VIS ICCV

[14] Tanks and Temples: Benchmarking Large-Scale Scene Reconstruction [J].

Knapitsch, Arno ;

Park, Jaesik ;

Zhou, Qian-Yi ;

Koltun, Vladlen .

ACM TRANSACTIONS ON GRAPHICS, 2017, 36 (04)

[15] MegaDepth: Learning Single-View Depth Prediction from Internet Photos [J].

Li, Zhengqi ;

Snavely, Noah .

2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :2041-2050

[16] A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation [J].

Mayer, Nikolaus ;

Ilg, Eddy ;

Hausser, Philip ;

Fischer, Philipp ;

Cremers, Daniel ;

Dosovitskiy, Alexey ;

Brox, Thomas .

2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :4040-4048

[17] Evaluation of large scale scene reconstruction [J].

Merrell, Paul ;

Mordohai, Philippos ;

Frahm, Jan-Michael ;

Pollefeys, Marc .

2007 IEEE 11TH INTERNATIONAL CONFERENCE ON COMPUTER VISION, VOLS 1-6, 2007, :3012-3019

[18]

Paschalidou Despoina, 2018, COMPUTER VISION PATT

[19] Playing for Data: Ground Truth from Computer Games [J].

Richter, Stephan R. ;

Vineet, Vibhav ;

Roth, Stefan ;

Koltun, Vladlen .

COMPUTER VISION - ECCV 2016, PT II, 2016, 9906 :102-118

[20]

Ros German, COMPUTER VISION PATT

← 1 2 3 4 →