Self-Supervised Learning of Neural Implicit Feature Fields for Camera Pose Refinement

被引:0
作者
Pietrantoni, Maxime [1 ,2 ]
Csurka, Gabriela [3 ]
Humenberger, Martin [3 ]
Sattler, Torsten [2 ]
机构
[1] Czech Tech Univ, Fac Elect Engn, Prague, Czech Republic
[2] Czech Tech Univ, Czech Inst Informat Robot & Cybernet, Prague, Czech Republic
[3] NAVER LABS Europe, Meylan, France
来源
2024 INTERNATIONAL CONFERENCE IN 3D VISION, 3DV 2024 | 2024年
关键词
RECOGNITION;
D O I
10.1109/3DV62453.2024.00139
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Visual localization techniques rely upon some underlying scene representation to localize against. These representations can be explicit such as 3D SFM map or implicit, such as a neural network that learns to encode the scene. The former requires sparse feature extractors and matchers to build the scene representation. The latter might lack geometric grounding not capturing the 3D structure of the scene well enough. This paper proposes to jointly learn the scene representation along with a 3D dense feature field and a 2D feature extractor whose outputs are embedded in the same metric space. Through a contrastive framework we align this volumetric field with the image-based extractor and regularize the latter with a ranking loss from learned surface information. We learn the underlying geometry of the scene with an implicit field through volumetric rendering and design our feature field to leverage intermediate geometric information encoded in the implicit field. The resulting features are discriminative and robust to viewpoint change while maintaining rich encoded information. Visual localization is then achieved by aligning the image-based features and the rendered volumetric features. We show the effectiveness of our approach on real-world scenes, demonstrating that our approach outperforms prior and concurrent work on leveraging implicit scene representations for localization.
引用
收藏
页码:484 / 494
页数:11
相关论文
共 86 条
[1]   Photometric Bundle Adjustment for Vision-Based SLAM [J].
Alismail, Hatem ;
Browning, Brett ;
Lucey, Simon .
COMPUTER VISION - ACCV 2016, PT IV, 2017, 10114 :324-341
[2]  
Arandjelovic R, 2018, IEEE T PATTERN ANAL, V40, P1437, DOI [10.1109/TPAMI.2017.2711011, 10.1109/CVPR.2016.572]
[3]   Wide Area Localization on Mobile Phones [J].
Arth, Clemens ;
Wagner, Daniel ;
Klopschitz, Manfred ;
Irschara, Arnold ;
Schmalstieg, Dieter .
2009 8TH IEEE INTERNATIONAL SYMPOSIUM ON MIXED AND AUGMENTED REALITY - SCIENCE AND TECHNOLOGY, 2009, :73-82
[4]   Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields [J].
Barron, Jonathan T. ;
Mildenhall, Ben ;
Verbin, Dor ;
Srinivasan, Pratul P. ;
Hedman, Peter .
2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, :19640-19648
[5]  
Bhayani Snehal, 2021, P IEEECVF INT C COMP, P5936
[6]   On the Limits of Pseudo Ground Truth in Visual Camera Re-localisation [J].
Brachmann, Eric ;
Humenberger, Martin ;
Rother, Carsten ;
Sattler, Torsten .
2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, :6198-6208
[7]   Visual Camera Re-Localization From RGB and RGB-D Images Using DSAC [J].
Brachmann, Eric ;
Rother, Carsten .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2022, 44 (09) :5847-5865
[8]   Expert Sample Consensus Applied to Camera Re-Localization [J].
Brachmann, Eric ;
Rother, Carsten .
2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, :7524-7533
[9]   Geometry-Aware Learning of Maps for Camera Localization [J].
Brahmbhatt, Samarth ;
Gu, Jinwei ;
Kim, Kihwan ;
Hays, James ;
Kautz, Jan .
2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :2616-2625
[10]   Emerging Properties in Self-Supervised Vision Transformers [J].
Caron, Mathilde ;
Touvron, Hugo ;
Misra, Ishan ;
Jegou, Herve ;
Mairal, Julien ;
Bojanowski, Piotr ;
Joulin, Armand .
2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, :9630-9640