Synthetic dataset generation for object-to-model deep learning in industrial applications

被引：29

作者：

Wong, Matthew Z. ^{[1
]}

Kunii, Kiyohito ^{[1
]}

Baylis, Max ^{[1
]}

Ong, Wai Hong ^{[1
]}

Kroupa, Pavel ^{[1
]}

Koller, Swen ^{[1
]}

机构：

[1] Imperial Coll London, Dept Comp, London, England

来源：

PEERJ COMPUTER SCIENCE | 2019年 / 2019卷 / 10期

关键词：

Industrial computer vision; Photogrammetry; Convolutional neural network; Computer science applications; 3D Modelling; Synthetic data; Deep learning with limited data;

D O I：

10.7717/peerj-cs.222

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

The availability of large image data sets has been a crucial factor in the success of deep learning-based classification and detection methods. Yet, while data sets for everyday objects are widely available, data for specific industrial use-cases (e.g., identifying packaged products in a warehouse) remains scarce. In such cases, the data sets have to be created from scratch, placing a crucial bottleneck on the deployment of deep learning techniques in industrial applications. We present work carried out in collaboration with a leading UK online supermarket, with the aim of creating a computer vision system capable of detecting and identifying unique supermarket products in a warehouse setting. To this end, we demonstrate a framework for using data synthesis to create an end-to-end deep learning pipeline, beginning with real-world objects and culminating in a trained model. Our method is based on the generation of a synthetic dataset from 3D models obtained by applying photogrammetry techniques to real-world objects. Using 100K synthetic images for 10 classes, an InceptionV3 convolutional neural network was trained, which achieved accuracy of 96% on a separately acquired test set of real supermarket product images. The image generation process supports automatic pixel annotation. This eliminates the prohibitively expensive manual annotation typically required for detection tasks. Based on this readily available data, a one-stage RetinaNet detector was trained on the synthetic, annotated images to produce a detector that can accurately localize and classify the specimen products in real-time.

引用

页数：18

共 22 条

[1]

Agisoft LLC, 2018, AG PHOT

[2]

Bergstra J, 2011, ADV NEURAL INFORM PR, P2546, DOI 10.5555/2986459.2986743

[3]

Bergstra J, 2012, J MACH LEARN RES, V13, P281

[4]

Chollet F., 2015, Keras

[5]

Deng J, 2009, PROC CVPR IEEE, P248, DOI 10.1109/CVPRW.2009.5206848

[6]

EyeCue Vision Technologies, 2018, QLON

[7]

Faugeras O. D., 1992, Computer Vision - ECCV '92. Second European Conference on Computer Vision Proceedings, P563

[8]

Girshick R., 2015, P IEEE INT C COMPUTE, DOI [DOI 10.1109/ICCV.2015.169, 10.1109/ICCV.2015.169]

[9]

He K., 2016, CVPR, DOI [10.1109/CVPR.2016.90, DOI 10.1109/CVPR.2016.90]

[10] Microsoft COCO: Common Objects in Context [J].

Lin, Tsung-Yi ;

Maire, Michael ;

Belongie, Serge ;

Hays, James ;

Perona, Pietro ;

Ramanan, Deva ;

Dollar, Piotr ;

Zitnick, C. Lawrence .

COMPUTER VISION - ECCV 2014, PT V, 2014, 8693 :740-755

← 1 2 3 →