A Machine Learning Approach for Layout Inference in Spreadsheets

被引:31
|
作者
Koci, Elvis [1 ]
Thiele, Maik [1 ]
Romero, Oscar [2 ]
Lehner, Wolfgang [1 ]
机构
[1] Tech Univ Dresden, Dept Comp Sci, Database Technol Grp, Dresden, Germany
[2] Univ Politecn Catalunya UPC BarcelonaTech, Dept Engn Serv & Sist Informacio, C-Jordi Girona 1,Compus Nord, Barcelona, Spain
来源
KDIR: PROCEEDINGS OF THE 8TH INTERNATIONAL JOINT CONFERENCE ON KNOWLEDGE DISCOVERY, KNOWLEDGE ENGINEERING AND KNOWLEDGE MANAGEMENT - VOL. 1 | 2016年
关键词
Speadsheets; Tabular; Layout; Structure; Machine Learning; Knowledge Discovery;
D O I
10.5220/0006052200770088
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Spreadsheet applications are one of the most used tools for content generation and presentation in industry and the Web. In spite of this success, there does not exist a comprehensive approach to automatically extract and reuse the richness of data maintained in this format. The biggest obstacle is the lack of awareness about the structure of the data in spreadsheets, which otherwise could provide the means to automatically understand and extract knowledge from these files. In this paper, we propose a classification approach to discover the layout of tables in spreadsheets. Therefore, we focus on the cell level, considering a wide range of features not covered before by related work. We evaluated the performance of our classifiers on a large dataset covering three different corpora from various domains. Finally, our work includes a novel technique for detecting and repairing incorrectly classified cells in a post-processing step. The experimental results show that our approach delivers very high accuracy bringing us a crucial step closer towards automatic table extraction.
引用
收藏
页码:77 / 88
页数:12
相关论文
共 50 条
  • [21] Machine learning transforms the inference of the nuclear equation of state
    Wang, Yongjia
    Li, Qingfeng
    FRONTIERS OF PHYSICS, 2023, 18 (06)
  • [22] Demystifying Membership Inference Attacks in Machine Learning as a Service
    Truex, Stacey
    Liu, Ling
    Gursoy, Mehmet Emre
    Yu, Lei
    Wei, Wenqi
    IEEE TRANSACTIONS ON SERVICES COMPUTING, 2021, 14 (06) : 2073 - 2089
  • [23] A NSGA-II-Based Layout Method for Cable Bundles With Branches Using Machine Learning
    Yang, Xu
    Zhou, Dejian
    Song, Wei
    IEEE ACCESS, 2021, 9 (09): : 90392 - 90401
  • [24] Machine learning transforms the inference of the nuclear equation of state
    Wang Yongjia
    Li Qingfeng
    Frontiers of Physics, 2023, 18 (06)
  • [25] GENECI: A novel evolutionary machine learning consensus-based approach for the inference of gene regulatory networks
    Segura-Ortiz, Adrian
    Garcia-Nieto, Jose
    Aldana-Montes, Jose F.
    Navas-Delgado, Ismael
    COMPUTERS IN BIOLOGY AND MEDICINE, 2023, 155
  • [26] Henna: Hierarchical Machine Learning Inference in Programmable Switches
    Tanyi-Jong Akem, Aristide
    Butun, Beyza
    Gucciardo, Michele
    Fiore, Marco
    PROCEEDINGS OF THE 1ST INTERNATIONAL WORKSHOP ON NATIVE NETWORK INTELLIGENCE, NATIVENI 2022, 2022, : 1 - 7
  • [27] MACHINE LEARNING-ENHANCED GENETIC ALGORITHM FOR ROBUST LAYOUT DESIGN IN DYNAMIC FACILITY LAYOUT PROBLEMS
    Amma, Vineetha Gopinathan Nair Radhamony
    Rasheedali, Shiyas Chekkot
    INTERNATIONAL JOURNAL OF INDUSTRIAL ENGINEERING-THEORY APPLICATIONS AND PRACTICE, 2023, 30 (06): : 1466 - 1485
  • [28] Deep Learning Approach to Biogeographical Ancestry Inference
    Qu, Yue
    Tran, Dat
    Ma, Wanli
    KNOWLEDGE-BASED AND INTELLIGENT INFORMATION & ENGINEERING SYSTEMS (KES 2019), 2019, 159 : 552 - 561
  • [29] Boosting Spectrum-Based Fault Localization for Spreadsheets with Product Metrics in a Learning Approach
    Mukhtar, Adil
    Hofer, Birgit
    Jannach, Dietmar
    Schekotihin, Konstantin
    Wotawa, Franz
    PROCEEDINGS OF THE 37TH IEEE/ACM INTERNATIONAL CONFERENCE ON AUTOMATED SOFTWARE ENGINEERING, ASE 2022, 2022,
  • [30] Near duplicate column identification: a machine learning approach
    Chevallier, Marc
    Boufares, Faouzi
    Grozavu, Nistor
    Rogovschi, Nicoleta
    Clairmont, Charly
    2021 IEEE SYMPOSIUM SERIES ON COMPUTATIONAL INTELLIGENCE (IEEE SSCI 2021), 2021,