Event-case correlation for process mining using probabilistic optimization

被引:8
作者
Bayomie, Dina [1 ,2 ]
Di Ciccio, Claudio [3 ]
Mendling, Jan [4 ]
机构
[1] Vienna Univ Econ & Business, Welthandelspl 1, A-1020 Vienna, Austria
[2] Cairo Univ, Gamaa St 1, Giza 12613, Egypt
[3] Sapienza Univ Rome, Viale Regina Elena 295, I-00161 Rome, Italy
[4] Humboldt Univ, Unter Linden 6, D-10099 Berlin, Germany
关键词
Process mining; Event correlation; Simulated annealing; Constraints; Association rules; PROCESS MODELS; CONFORMANCE CHECKING; AUTOMATED DISCOVERY; ANALYTICS; ACCURATE;
D O I
10.1016/j.is.2023.102167
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Process mining supports the analysis of the actual behavior and performance of business processes using event logs. An essential requirement is that every event in the log must be associated with a unique case identifier (e.g., the order ID of an order-to-cash process). In reality, however, this case identifier may not always be present, especially when logs are acquired from different systems or extracted from non-process-aware information systems. In such settings, the event log needs to be pre-processed by grouping events into cases - an operation known as event correlation. Existing techniques for correlating events have worked with assumptions to make the problem tractable: some assume the generative processes to be acyclic, while others require heuristic information or user input. Moreover, they abstract the log to activities and timestamps, and miss the opportunity to use data attributes. In this paper, we lift these assumptions and propose a new technique called EC-SA-Data based on probabilistic optimization. The technique takes as inputs a sequence of timestamped events (the log without case IDs), a process model describing the underlying business process, and constraints over the event attributes. Our approach returns an event log in which every event is associated with a case identifier. The technique allows users to flexibly incorporate rules on process knowledge and data constraints. The approach minimizes the misalignment between the generated log and the input process model, maximizes the support of the given data constraints over the correlated log, and the variance between activity durations across cases. Our experiments with various real-life datasets show the advantages of our approach over the state of the art.(c) 2023 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
引用
收藏
页数:28
相关论文
共 81 条
[1]  
Abiteboul S., 2003, Proceedings of the 29th international conference on Very large data bases-Volume 29, VLDB '2003, P668, DOI 10.1016/B978-012722442-8/50065-3
[2]  
Achtert E., 2008, Stat Anal Data Mining: The ASA Data Sci J, V1, P111, DOI [DOI 10.1002/SAM.10012, DOI 10.1002/sam.10012, 10.1002/sam.10012]
[3]   Measuring precision of modeled behavior [J].
Adriansyah, A. ;
Munoz-Gama, J. ;
Carmona, J. ;
van Dongen, B. F. ;
van der Aalst, W. M. P. .
INFORMATION SYSTEMS AND E-BUSINESS MANAGEMENT, 2015, 13 (01) :37-67
[4]   Conformance Checking using Cost-Based Fitness Analysis [J].
Adriansyah, A. ;
van Dongen, B. F. ;
van der Aalst, W. M. P. .
15TH IEEE INTERNATIONAL ENTERPRISE DISTRIBUTED OBJECT COMPUTING CONFERENCE (EDOC 2011), 2011, :55-64
[5]   Toward an Automated Labeling of Event Log Attributes [J].
Andaloussi, Amine Abbad ;
Burattin, Andrea ;
Weber, Barbara .
ENTERPRISE, BUSINESS-PROCESS AND INFORMATION SYSTEMS MODELING, BPMDS 2018 AND EMMSAD 2018, 2018, 318 :82-96
[6]   Optimal exact experimental designs with correlated, errors through a simulated annealing algorithm [J].
Angelis, L ;
Bora-Senta, E ;
Moyssiadis, C .
COMPUTATIONAL STATISTICS & DATA ANALYSIS, 2001, 37 (03) :275-296
[7]  
[Anonymous], 2012, P C MOVE MEANINGFUL, DOI DOI 10.1007/978-3-642-33606-5_19
[8]   An empirical comparison of Tabu Search, Simulated Annealing, and Genetic Algorithms for facilities location problems [J].
Arostegui, Marvin A., Jr. ;
Kadipasaoglu, Sukran N. ;
Khumawala, Basheer M. .
INTERNATIONAL JOURNAL OF PRODUCTION ECONOMICS, 2006, 103 (02) :742-754
[9]  
Askarzadeh A, 2016, IEEE SYS MAN CYBERN, P4626, DOI 10.1109/SMC.2016.7844961
[10]   The connection between process complexity of event sequences and models discovered by process mining [J].
Augusto, Adriano ;
Mendling, Jan ;
Vidgof, Maxim ;
Wurm, Bastian .
INFORMATION SCIENCES, 2022, 598 :196-215