Incremental Instance-Oriented 3D Semantic Mapping via RGB-D Cameras for Unknown Indoor Scene

被引：10

作者：

Li, Wei ^{[1
]}

Gu, Junhua ^{[2
]}

Chen, Benwen ^{[2
]}

Han, Jungong ^{[3
]}

机构：

[1] Hebei Univ Technol, State Key Lab Reliabil & Intelligence Elect Equip, Key Lab Electromagnet Field & Elect Apparat Relia, Sch Elect Engn, Tianjin 300401, Peoples R China

[2] Hebei Univ Technol, Sch Artificial Intelligence, Key Lab Big Data Comp, Tianjin 300401, Peoples R China

[3] Univ Warwick, WMG Data Sci, Coventry CV4 7AL, W Midlands, England

来源：

DISCRETE DYNAMICS IN NATURE AND SOCIETY | 2020年 / 2020卷

关键词：

SALIENCY DETECTION; RECOGNITION;

D O I：

10.1155/2020/2528954

中图分类号：

O1 [数学];

学科分类号：

0701 ; 070101 ;

摘要：

Scene parsing plays a crucial role when accomplishing human-robot interaction tasks. As the "eye" of the robot, RGB-D camera is one of the most important components for collecting multiview images to construct instance-oriented 3D environment semantic maps, especially in unknown indoor scenes. Although there are plenty of studies developing accurate object-level mapping systems with different types of cameras, these methods either process the instance segmentation problem in completed mapping or suffer from a critical real-time issue due to heavy computation processing required. In this paper, we propose a novel method to incrementally build instance-oriented 3D semantic maps directly from images acquired by the RGB-D camera. To ensure an efficient reconstruction of 3D objects with semantic and instance IDs, the input RGB images are operated by a real-time deep-learned object detector. To obtain accurate point cloud cluster, we adopt the Gaussian mixture model as an optimizer after processing 2D to 3D projection. Next, we present a data association strategy to update class probabilities across the frames. Finally, a map integration strategy fuses information about their 3D shapes, locations, and instance IDs in a faster way. We evaluate our system on different indoor scenes including offices, bedrooms, and living rooms from the SceneNN dataset, and the results show that our method not only builds the instance-oriented semantic map efficiently but also enhances the accuracy of the individual instance in the scene.

引用

页数：10

共 47 条

[1] Tutorial Point Cloud Library Three-Dimensional Object Recognition and 6 DOF Pose Estimation [J].

Aldoma, Aitor ;

Marton, Zoltan-Csaba ;

Tombari, Federico ;

Wohlkinger, Walter ;

Potthast, Christian ;

Zeisl, Bernhard ;

Rusu, Radu Bogdan ;

Gedikli, Suat ;

Vincze, Markus .

IEEE ROBOTICS & AUTOMATION MAGAZINE, 2012, 19 (03) :80-91

[2]

[Anonymous], 2017, JOINT 2D 3D SEMANTIC

[3]

[Anonymous], 2012, MIT CSAIL TR

[4]

[Anonymous], SEMIDENSE 3D SEMANTI

[5]

[Anonymous], 2007, LECT NOTES COMPUT SC, DOI DOI 10.1007/S11263-014-0733-5

[6]

[Anonymous], P ROB SCI SYST 13 JU

[7]

Civera J., 2011, STRUCTURE MOTION USI

[8] ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes [J].

Dai, Angela ;

Chang, Angel X. ;

Savva, Manolis ;

Halber, Maciej ;

Funkhouser, Thomas ;

Niessner, Matthias .

30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017), 2017, :2432-2443

[9] BundleFusion: Real-Time Globally Consistent 3D Reconstruction Using On-the-Fly Surface Reintegration [J].

Dai, Angela ;

Niessner, Matthias ;

Zollhofer, Michael ;

Izadi, Shahram ;

Theobalt, Christian .

ACM TRANSACTIONS ON GRAPHICS, 2017, 36 (03)

[10]

Eckart B, 2013, IEEE INT C INT ROBOT, P4355, DOI 10.1109/IROS.2013.6696981

← 1 2 3 4 5 →