Depth from Camera Motion and Object Detection

被引：16

作者：

Griffin, Brent A. ^{[1
]}

Corso, Jason J. ^{[2
]}

机构：

[1] Univ Michigan, Ann Arbor, MI 48109 USA

[2] Stevens Inst Artificial Intelligence, Hoboken, NJ USA

来源：

2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021 | 2021年

关键词：

PARALLAX; SIZE; CUE;

D O I：

10.1109/CVPR46437.2021.00145

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

This paper addresses the problem of learning to estimate the depth of detected objects given some measurement of camera motion (e.g., from robot kinematics or vehicle odometry). We achieve this by 1) designing a recurrent neural network (DBox) that estimates the depth of objects using a generalized representation of bounding boxes and uncalibrated camera movement and 2) introducing the Object Depth via Motion and Detection Dataset (ODMD). ODMD training data are extensible and configurable, and the ODMD benchmark includes 21,600 examples across four validation and test sets. These sets include mobile robot experiments using an end-effector camera to locate objects from the YCB dataset and examples with perturbations added to camera motion or bounding box data. In addition to the ODMD benchmark, we evaluate DBox in other monocular application domains, achieving state-of-the-art results on existing driving and robotics benchmarks and estimating the depth of objects using a camera phone.

引用

页码：1397 / 1406

页数：10

共 55 条

[1]

Brachmann Eric, 2016, IEEE C COMP VIS PATT

[2] Cascade R-CNN: Delving into High Quality Object Detection [J].

Cai, Zhaowei ;

Vasconcelos, Nuno .

2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :6154-6162

[3] Benchmarking in Manipulation Research Using the Yale-CMU-Berkeley Object and Model Set [J].

Calli, Berk ;

Walsman, Aaron ;

Singh, Arjun ;

Srinivasa, Siddhartha ;

Abbeel, Pieter ;

Dollar, Aaron M. .

IEEE ROBOTICS & AUTOMATION MAGAZINE, 2015, 22 (03) :36-52

[4]

Carion Nicolas, 2020, EUROPEAN C COMPUTER

[5] The Cityscapes Dataset for Semantic Urban Scene Understanding [J].

Cordts, Marius ;

Omran, Mohamed ;

Ramos, Sebastian ;

Rehfeld, Timo ;

Enzweiler, Markus ;

Benenson, Rodrigo ;

Franke, Uwe ;

Roth, Stefan ;

Schiele, Bernt .

2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :3213-3223

[6]

Deng J, 2009, PROC CVPR IEEE, P248, DOI 10.1109/CVPRW.2009.5206848

[7]

Deng XK, 2020, IEEE INT CONF ROBOT, P3665, DOI [10.1109/ICRA40945.2020.9196714, 10.1109/icra40945.2020.9196714]

[8] A 2D-3D Object Detection System for Updating Building Information Models with Mobile Robots [J].

Ferguson, Max ;

Law, Kincho .

2019 IEEE WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV), 2019, :1357-1365

[9] MOTION PARALLAX AND ABSOLUTE DISTANCE [J].

FERRIS, SH .

JOURNAL OF EXPERIMENTAL PSYCHOLOGY, 1972, 95 (02) :258-&

[10] Bayesian Spatial Kernel Smoothing for Scalable Dense Semantic Mapping [J].

Gan, Lu ;

Zhang, Ray ;

Grizzle, Jessy W. ;

Eustice, Ryan M. ;

Ghaffari, Maani .

IEEE ROBOTICS AND AUTOMATION LETTERS, 2020, 5 (02) :790-797

← 1 2 3 4 5 6 →