A Machine Learning Approach for Layout Inference in Spreadsheets

被引:32
作者
Koci, Elvis [1 ]
Thiele, Maik [1 ]
Romero, Oscar [2 ]
Lehner, Wolfgang [1 ]
机构
[1] Tech Univ Dresden, Dept Comp Sci, Database Technol Grp, Dresden, Germany
[2] Univ Politecn Catalunya UPC BarcelonaTech, Dept Engn Serv & Sist Informacio, C-Jordi Girona 1,Compus Nord, Barcelona, Spain
来源
KDIR: PROCEEDINGS OF THE 8TH INTERNATIONAL JOINT CONFERENCE ON KNOWLEDGE DISCOVERY, KNOWLEDGE ENGINEERING AND KNOWLEDGE MANAGEMENT - VOL. 1 | 2016年
关键词
Speadsheets; Tabular; Layout; Structure; Machine Learning; Knowledge Discovery;
D O I
10.5220/0006052200770088
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Spreadsheet applications are one of the most used tools for content generation and presentation in industry and the Web. In spite of this success, there does not exist a comprehensive approach to automatically extract and reuse the richness of data maintained in this format. The biggest obstacle is the lack of awareness about the structure of the data in spreadsheets, which otherwise could provide the means to automatically understand and extract knowledge from these files. In this paper, we propose a classification approach to discover the layout of tables in spreadsheets. Therefore, we focus on the cell level, considering a wide range of features not covered before by related work. We evaluated the performance of our classifiers on a large dataset covering three different corpora from various domains. Finally, our work includes a novel technique for detecting and repairing incorrectly classified cells in a post-processing step. The experimental results show that our approach delivers very high accuracy bringing us a crucial step closer towards automatic table extraction.
引用
收藏
页码:77 / 88
页数:12
相关论文
共 17 条
[1]   Header and unit inference for spreadsheets through spatial analyses [J].
Abraham, R ;
Erwig, M .
2004 IEEE SYMPOSIUM ON VISUAL LANGUAGES AND HUMAN CENTRIC COMPUTING: PROCEEDINGS, 2004, :165-172
[2]   Schema Extraction for Tabular Data on the Web [J].
Adelfio, Marco D. ;
Samet, Hanan .
PROCEEDINGS OF THE VLDB ENDOWMENT, 2013, 6 (06) :421-432
[3]  
Barik T., 2015, MSR 15
[4]   SmcHD1, containing a structural-maintenance-of-chromosomes hinge domain, has a critical role in X inactivation [J].
Blewitt, Marnie E. ;
Gendrel, Anne-Valerie ;
Pang, Zhenyi ;
Sparrow, Duncan B. ;
Whitelaw, Nadia ;
Craig, Jeffrey M. ;
Apedaile, Anwyn ;
Hilton, Douglas J. ;
Dunwoodie, Sally L. ;
Brockdorff, Neil ;
Kay, Graham F. ;
Whitelaw, Emma .
NATURE GENETICS, 2008, 40 (05) :663-669
[5]   Random forests [J].
Breiman, L .
MACHINE LEARNING, 2001, 45 (01) :5-32
[6]  
Chen Z., 2013, P INT WORKSH SEM SEA, DOI DOI 10.1145/2509908.2509909
[7]   Integrating Spreadsheet Data via Accurate and Low-Effort Extraction [J].
Chen, Zhe ;
Cafarella, Michael .
PROCEEDINGS OF THE 20TH ACM SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING (KDD'14), 2014, :1126-1135
[8]  
Crestan E, 2011, P 4 ACM INT C WEB SE, P545, DOI DOI 10.1145/1935826.1935904
[9]  
Eberius J., 2015, BDC 15
[10]   DeExcelerator: A Framework for Extracting Relational Data From Partially Structured Documents [J].
Eberius, Julian ;
Werner, Christoper ;
Thiele, Maik ;
Braunschweig, Katrin ;
Dannecker, Lars ;
Lehner, Wolfgang .
PROCEEDINGS OF THE 22ND ACM INTERNATIONAL CONFERENCE ON INFORMATION & KNOWLEDGE MANAGEMENT (CIKM'13), 2013,