Fast Object Segmentation Learning with Kernel-based Methods for Robotics

被引：4

作者：

Ceola, Federico ^{[1
,2
,3
]}

Maiettini, Elisa ^{[1
]}

Pasquale, Giulia ^{[1
]}

Rosasco, Lorenzo ^{[2
,3
,4
,5
]}

Natale, Lorenzo ^{[1
]}

机构：

[1] Ist Italiano Tecnol, Humanoid Sensing & Percept, Genoa, Italy

[2] Univ Genoa, Lab Computat & Stat Learning, Genoa, Italy

[3] Univ Genoa, Dipartimento Informat Bioingn Robot & Ingn Sistem, Genoa, Italy

[4] Ist Italiano Tecnol, Genoa, Italy

[5] MIT, 77 Massachusetts Ave, Cambridge, MA 02139 USA

来源：

2021 IEEE INTERNATIONAL CONFERENCE ON ROBOTICS AND AUTOMATION (ICRA 2021) | 2021年

基金：

英国工程与自然科学研究理事会;

关键词：

D O I：

10.1109/ICRA48506.2021.9561758

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Object segmentation is a key component in the visual system of a robot that performs tasks like grasping and object manipulation, especially in presence of occlusions. Like many other computer vision tasks, the adoption of deep architectures has made available algorithms that perform this task with remarkable performance. However, adoption of such algorithms in robotics is hampered by the fact that training requires large amount of computing time and it cannot be performed on-line. In this work, we propose a novel architecture for object segmentation, that overcomes this problem and provides comparable performance in a fraction of the time required by the state-of-the-art methods. Our approach is based on a pre-trained Mask R-CNN, in which various layers have been replaced with a set of classifiers and regressors that are retrained for a new task. We employ an efficient Kernel-based method that allows for fast training on large scale problems. Our approach is validated on the YCB-Video dataset which is widely adopted in the computer vision and robotics community, demonstrating that we can achieve and even surpass performance of the state-of-the-art, with a significant reduction (similar to 6 x) of the training time. The code to reproduce the experiments is publicly available on GitHub(1).

引用

页码：13581 / 13588

页数：8

共 44 条

[1] Deep Watershed Transform for Instance Segmentation [J].

Bai, Min ;

Urtasun, Raquel .

30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017), 2017, :2858-2866

[2]

Caelles Sergi, 2019, arXiv:1905.00737

[3]

Ceola F., 2020, ARXIV201112790

[4] The devil is in the details: an evaluation of recent feature encoding methods [J].

Chatfield, Ken ;

Lempitsky, Victor ;

Vedaldi, Andrea ;

Zisserman, Andrew .

PROCEEDINGS OF THE BRITISH MACHINE VISION CONFERENCE 2011, 2011,

[5] Instance-aware Semantic Segmentation via Multi-task Network Cascades [J].

Dai, Jifeng ;

He, Kaiming ;

Sun, Jian .

2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :3150-3158

[6]

Deng XK, 2020, IEEE INT CONF ROBOT, P3665, DOI [10.1109/ICRA40945.2020.9196714, 10.1109/icra40945.2020.9196714]

[7]

Denninger Maximilian, 2020, RSS 2020

[8] The Pascal Visual Object Classes (VOC) Challenge [J].

Everingham, Mark ;

Van Gool, Luc ;

Williams, Christopher K. I. ;

Winn, John ;

Zisserman, Andrew .

INTERNATIONAL JOURNAL OF COMPUTER VISION, 2010, 88 (02) :303-338

[9] Learning Hierarchical Features for Scene Labeling [J].

Farabet, Clement ;

Couprie, Camille ;

Najman, Laurent ;

LeCun, Yann .

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2013, 35 (08) :1915-1929

[10]

Fulkerson B, 2009, IEEE I CONF COMP VIS, P670

← 1 2 3 4 5 →