Multi-label Annotation for Visual Multi-Task Learning Models

被引：0

作者：

Sharma, Gaurang ^{[1
]}

Angleraud, Alexandre ^{[1
]}

Pieters, Roel ^{[1
]}

机构：

[1] Tampere Univ, Tampere, Finland

来源：

2023 SEVENTH IEEE INTERNATIONAL CONFERENCE ON ROBOTIC COMPUTING, IRC 2023 | 2023年

关键词：

Multi-task models; image annotation and augmentation; object detection and segmentation; keypoint detection;

D O I：

10.1109/IRC59093.2023.00012

中图分类号：

TP301 [理论、方法];

学科分类号：

081202 ;

摘要：

Deep learning requires large amounts of data, and a well-defined pipeline for labeling and augmentation. Current solutions support numerous computer vision tasks with dedicated annotation types and formats, such as bounding boxes, polygons, and key points. These annotations can be combined into a single data format to benefit approaches such as multi-task models. However, to our knowledge, no available labeling tool supports the export functionality for a combined benchmark format, and no augmentation library supports transformations for the combination of all. In this work, these functionalities are presented, with visual data annotation and augmentation to train a multi-task model (object detection, segmentation, and key point extraction). The tools are demonstrated in two robot perception use cases. For more details, please visit https://gaurangsharma18.github.io/website/MultiLabel/index.html.

引用

页码：31 / 34

页数：4

共 29 条

[1] Abadi M., 2015, TensorFlow. Large-Scale Machine Learning on Heterogeneous Systems, V1, DOI DOI 10.5281/ZENODO.8071599
[2] Albumentations, 2023, US
[3] Image annotation: Then and now
Bhagat, P. K.
Choudhary, P.
[J]. IMAGE AND VISION COMPUTING, 2018, 80 : 1 - 23
[4] Bradski G, 2000, DR DOBBS J, V25, P120
[5] Rethinking interactive image segmentation: Feature space annotation
Bragantini, Jordao
Falcao, Alexandre X.
Najman, Laurent
[J]. PATTERN RECOGNITION, 2022, 131
[6] Detectron2, 2023, ABOUT US
[7] Local keypoint-based Faster R-CNN
Ding, Xintao
Li, Qingde
Cheng, Yongqiang
Wang, Jinbao
Bian, Weixin
Jie, Biao
[J]. APPLIED INTELLIGENCE, 2020, 50 (10) : 3007 - 3022
[8] The Pascal Visual Object Classes (VOC) Challenge
Everingham, Mark
Van Gool, Luc
Williams, Christopher K. I.
Winn, John
Zisserman, Andrew
[J]. INTERNATIONAL JOURNAL OF COMPUTER VISION, 2010, 88 (02) : 303 - 338
[9] Fifty C, 2021, ADV NEUR IN, V34
[10] Harris C., 1988, ALVEY VISION C, V15, P10, DOI [10.5244/c.2.23, DOI 10.5244/C.2.23]

← 1 2 3 →