Scalable graph-based OLAP analytics over process execution data

被引:0
作者
Seyed-Mehdi-Reza Beheshti
Boualem Benatallah
Hamid Reza Motahari-Nezhad
机构
[1] University of New South Wales,School of Computer Science and Engineering
[2] IBM Almaden Research Center,undefined
来源
Distributed and Parallel Databases | 2016年 / 34卷
关键词
Process analytics; Business analytics; Bigdata analytics; Graph OLAP; OLAP;
D O I
暂无
中图分类号
学科分类号
摘要
In today’s knowledge-, service-, and cloud-based economy, businesses accumulate massive amounts of data from a variety of sources. In order to understand businesses one may need to perform considerable analytics over large hybrid collections of heterogeneous and partially unstructured data that is captured related to the process execution. This data, usually modeled as graphs, increasingly come to show all the typical properties of big data: wide physical distribution, diversity of formats, non-standard data models, independently-managed and heterogeneous semantics. We use the term big process graph to refer to such large hybrid collections of heterogeneous and partially unstructured process related execution data. Online analytical processing (OLAP) of big process graph is challenging as the extension of existing OLAP techniques to analysis of graphs is not straightforward. Moreover, process data analysis methods should be capable of processing and querying large amount of data effectively and efficiently, and therefore have to be able to scale well with the infrastructure’s scale. While traditional analytics solutions (relational DBs, data warehouses and OLAP), do a great job in collecting data and providing answers on known questions, key business insights remain hidden in the interactions among objects: it will be hard to discover concept hierarchies for entities based on both data objects and their interactions in process graphs. In this paper, we introduce a framework and a set of methods to support scalable graph-based OLAP analytics over process execution data. The goal is to facilitate the analytics over big process graph through summarizing the process graph and providing multiple views at different granularity. To achieve this goal, we present a model for process OLAP (P-OLAP) and define OLAP specific abstractions in process context such as process cubes, dimensions, and cells. We present a MapReduce-based graph processing engine, to support big data analytics over process graphs. We have implemented the P-OLAP framework and integrated it into our existing process data analytics platform, ProcessAtlas, which introduces a scalable architecture for querying, exploration and analysis of large process data. We report on experiments performed on both synthetic and real-world datasets that show the viability and efficiency of the approach.
引用
收藏
页码:379 / 423
页数:44
相关论文
共 63 条
[1]  
Aalst WMPVD(2003)Workflow mining: a survey of issues and approaches Data Knowl. Eng. 47 237-267
[2]  
Dongen BFV(2012)Service mining: using process mining to discover, check, and improve service behavior IEEE Trans. Serv. Comput. 99 1-73
[3]  
Herbst J(2009)Extending SPARQL with regular expression patterns (for querying RDF) J. Web Sem. 7 57-88
[4]  
Maruster L(2014)Representation and querying of unfair evaluations in social rating systems Comput. Secur. 41 68-71
[5]  
Schimm G(2003)Intelligent business analytics: a tool to build decision-support systems for ebusinesses BT Technol. J. 21 65-22
[6]  
Weijters AJMM(2009)Linked data-the story so far Int. J. Semant. Web Inf. Syst. 5 1-74
[7]  
Aalst WMPVD(1997)An overview of data warehousing and OLAP technology SIGMOD Rec. 26 65-98
[8]  
Alkhateeb F(2011)An overview of business intelligence technology Commun. ACM 54 88-1000
[9]  
Baget JF(2009)Semantics preserving SPARQL-to-SQL translation Data Knowl. Eng. 68 973-865
[10]  
Euzenat J(2010)RDFProv: a relational RDF store for querying and managing scientific workflow provenance Data Knowl. Eng. 69 836-113