Catalog Integration of Heterogeneous and Volatile Product Data

被引:0
作者
Schmidts, Oliver [1 ]
Kraft, Bodo [1 ]
Winkens, Marvin [1 ]
Zuendorf, Albert [2 ]
机构
[1] Univ Appl Sci, FH Aachen, Julich, Germany
[2] Univ Kassel, Kassel, Germany
来源
DATA MANAGEMENT TECHNOLOGIES AND APPLICATIONS, DATA 2020 | 2021年 / 1446卷
关键词
Catalog integration; Data integration; Data quality; Label prediction; Machine learning; Neural network applications;
D O I
10.1007/978-3-030-83014-4_7
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The integration of frequently changing, volatile product data from different manufacturers into a single catalog is a significant challenge for small and medium-sized e-commerce companies. They rely on timely integrating product data to present them aggregated in an online shop without knowing format specifications, concept understanding of manufacturers, and data quality. Furthermore, format, concepts, and data quality may change at any time. Consequently, integrating product catalogs into a single standardized catalog is often a laborious manual task. Current strategies to streamline or automate catalog integration use techniques based on machine learning, word vectorization, or semantic similarity. However, most approaches struggle with low-quality or real-world data. We propose Attribute Label Ranking (ALR) as a recommendation engine to simplify the integration process of previously unknown, proprietary tabular format into a standardized catalog for practitioners. We evaluate ALR by focusing on the impact of different neural network architectures, language features, and semantic similarity. Additionally, we consider metrics for industrial application and present the impact of ALR in production and its limitations.
引用
收藏
页码:134 / 153
页数:20
相关论文
共 24 条
[1]  
Allweyer O., 2020, DATA 2020, P67
[2]  
[Anonymous], 2005, P 2005 ACM SIGMOD IN, DOI DOI 10.1145/1066157.1066283
[3]  
Bernstein PA, 2011, PROC VLDB ENDOW, V4, P695
[4]   Using the Semantic Web as a Source of Training Data [J].
Bizer, Christian ;
Primpeli, Anna ;
Peeters, Ralph .
Datenbank-Spektrum, 2019, 19 (02) :127-135
[5]  
Bojanowski P, 2017, Arxiv, DOI arXiv:1607.04606
[6]   Generating Schema Labels through Dataset Content Analysis [J].
Chen, Zhiyu ;
Jia, Haiyan ;
Heflin, Jeff ;
Davison, Brian D. .
COMPANION PROCEEDINGS OF THE WORLD WIDE WEB CONFERENCE 2018 (WWW 2018), 2018, :1515-1522
[7]  
Comito C., 2006, 11 IEEE S COMP COMM, P88, DOI [10.1109/ISCC.2006.19, DOI 10.1109/ISCC.2006.19]
[8]   An evolutionary approach to complex schema matching [J].
de Carvalho, Moises Gomes ;
Laender, Alberto H. F. ;
Goncalves, Marcos Andre ;
da Silva, Altigran S. .
INFORMATION SYSTEMS, 2013, 38 (03) :302-316
[9]   Orchid:: Integrating schema mapping and ETL [J].
Dessloch, Stefan ;
Hernandez, Mauricio A. ;
Wisnesky, Ryan ;
Radwan, Ahmed ;
Zhou, Jindan .
2008 IEEE 24TH INTERNATIONAL CONFERENCE ON DATA ENGINEERING, VOLS 1-3, 2008, :1307-+
[10]  
Devlin J, 2019, 2019 CONFERENCE OF THE NORTH AMERICAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS: HUMAN LANGUAGE TECHNOLOGIES (NAACL HLT 2019), VOL. 1, P4171