A Machine Learning Approach for Layout Inference in Spreadsheets

被引:31
|
作者
Koci, Elvis [1 ]
Thiele, Maik [1 ]
Romero, Oscar [2 ]
Lehner, Wolfgang [1 ]
机构
[1] Tech Univ Dresden, Dept Comp Sci, Database Technol Grp, Dresden, Germany
[2] Univ Politecn Catalunya UPC BarcelonaTech, Dept Engn Serv & Sist Informacio, C-Jordi Girona 1,Compus Nord, Barcelona, Spain
来源
KDIR: PROCEEDINGS OF THE 8TH INTERNATIONAL JOINT CONFERENCE ON KNOWLEDGE DISCOVERY, KNOWLEDGE ENGINEERING AND KNOWLEDGE MANAGEMENT - VOL. 1 | 2016年
关键词
Speadsheets; Tabular; Layout; Structure; Machine Learning; Knowledge Discovery;
D O I
10.5220/0006052200770088
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Spreadsheet applications are one of the most used tools for content generation and presentation in industry and the Web. In spite of this success, there does not exist a comprehensive approach to automatically extract and reuse the richness of data maintained in this format. The biggest obstacle is the lack of awareness about the structure of the data in spreadsheets, which otherwise could provide the means to automatically understand and extract knowledge from these files. In this paper, we propose a classification approach to discover the layout of tables in spreadsheets. Therefore, we focus on the cell level, considering a wide range of features not covered before by related work. We evaluated the performance of our classifiers on a large dataset covering three different corpora from various domains. Finally, our work includes a novel technique for detecting and repairing incorrectly classified cells in a post-processing step. The experimental results show that our approach delivers very high accuracy bringing us a crucial step closer towards automatic table extraction.
引用
收藏
页码:77 / 88
页数:12
相关论文
共 50 条
  • [1] A Study on Machine Learning Web Service Using Spreadsheets
    Yoon, Chi-Yurl
    Kang, Shin-Gak
    2017 INTERNATIONAL CONFERENCE ON INFORMATION AND COMMUNICATION TECHNOLOGY CONVERGENCE (ICTC), 2017, : 760 - 765
  • [2] A three-stage machine learning and inference approach for educational data
    Da, Ting
    SCIENTIFIC REPORTS, 2025, 15 (01):
  • [3] Machine learning in causal inference for epidemiology
    Moccia, Chiara
    Moirano, Giovenale
    Popovic, Maja
    Pizzi, Costanza
    Fariselli, Piero
    Richiardi, Lorenzo
    Ekstrom, Claus Thorn
    Maule, Milena
    EUROPEAN JOURNAL OF EPIDEMIOLOGY, 2024, 39 (10) : 1097 - 1108
  • [4] Video QoE Inference with Machine Learning
    Tisa-Selma
    Bentaleb, Abdelhak
    Harous, Saad
    IWCMC 2021: 2021 17TH INTERNATIONAL WIRELESS COMMUNICATIONS & MOBILE COMPUTING CONFERENCE (IWCMC), 2021, : 1048 - 1053
  • [5] What Would a Graph Look Like in This Layout? A Machine Learning Approach to Large Graph Visualization
    Kwon, Oh-Hyun
    Crnovrsanin, Tarik
    Ma, Kwan-Liu
    IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, 2018, 24 (01) : 478 - 488
  • [6] A machine learning approach for modeling irregular regions with multiple owners in wind farm layout design
    Reddy, Sohail R.
    ENERGY, 2021, 220
  • [7] Bibliometric Study on the Use of Machine Learning as Resolution Technique for Facility Layout Problems
    Burggraef, Peter
    Wagner, Johannes
    Heinbach, Benjamin
    IEEE ACCESS, 2021, 9 : 22569 - 22586
  • [8] Runtime Data Layout Scheduling for Machine Learning Dataset
    You, Yang
    Demmel, James
    2017 46TH INTERNATIONAL CONFERENCE ON PARALLEL PROCESSING (ICPP), 2017, : 452 - 461
  • [9] Linear kitchen layout design via machine learning
    Pejic, Jelena
    Pejic, Petar
    AI EDAM-ARTIFICIAL INTELLIGENCE FOR ENGINEERING DESIGN ANALYSIS AND MANUFACTURING, 2022, 36
  • [10] Recent Developments in Causal Inference and Machine Learning
    Brand, Jennie E.
    Zhou, Xiang
    Xie, Yu
    ANNUAL REVIEW OF SOCIOLOGY, 2023, 49 : 81 - 110