<p>Event logs are data sets recording the executions (called cases) of a business process. Several process discovery algorithms have been defined to mine event logs and discover models of how activities of logged processes are being executed (activity traces). In the process discovery problems, the Pareto principle plays an important role. In fact, it is quite common that a large portion of log traces is held by a small fraction of top-frequent variants. Hence, accounting for the expected Pareto distribution of event logs, traditional process discovery algorithms are commonly used to discover process models by analyzing the prevalent trace behaviors. However, the Pareto principle is not always verified, especially in complex processes, where the majority of traces in the event log is often spanned on a high number of top-frequent trace-variants. In addition, traditional process discovery algorithms perform an offline analysis of event logs, assuming that logged processes remain in a steady state over time. But, the steady state is rarely the real-world case due to conceptual drifts. In this study, we use the traditional process discovery algorithms under the dynamic conditions of real-world processes. To this aim, we define two approaches, namely <Emphasis FontCategory="SansSerif">FAIRY</Emphasis><InlineEquation ID="IEq1"> <EquationSource Format="TEX">\( ^{\textsf{P}} \)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow /> <mi mathvariant="sans-serif">P</mi> </mmultiscripts> </math></EquationSource> </InlineEquation> and <Emphasis FontCategory="SansSerif">FAIRY</Emphasis><InlineEquation ID="IEq2"> <EquationSource Format="TEX">\( ^{\textsf{NP}} \)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow /> <mi mathvariant="sans-serif">NP</mi> </mmultiscripts> </math></EquationSource> </InlineEquation>, which detect drifts in the conformance of traces to process models and discover new process models on drifted traces. In <Emphasis FontCategory="SansSerif">FAIRY</Emphasis><InlineEquation ID="IEq3"> <EquationSource Format="TEX">\( ^{\textsf{P}} \)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow /> <mi mathvariant="sans-serif">P</mi> </mmultiscripts> </math></EquationSource> </InlineEquation>, a new process model is discovered on an extraction-based representation of a drift. In <Emphasis FontCategory="SansSerif">FAIRY</Emphasis><InlineEquation ID="IEq4"> <EquationSource Format="TEX">\( ^{\textsf{NP}} \)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow /> <mi mathvariant="sans-serif">NP</mi> </mmultiscripts> </math></EquationSource> </InlineEquation>, a new process model is discovered on an abstraction-based representation of a drift. The experimental results analyze the performance of the proposed approaches, also compared to a few related methods, showing the effectiveness of <Emphasis FontCategory="SansSerif">FAIRY</Emphasis><InlineEquation ID="IEq5"> <EquationSource Format="TEX">\( ^{\textsf{P}} \)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow /> <mi mathvariant="sans-serif">P</mi> </mmultiscripts> </math></EquationSource> </InlineEquation> in Pareto cases and <Emphasis FontCategory="SansSerif">FAIRY</Emphasis><InlineEquation ID="IEq6"> <EquationSource Format="TEX">\( ^{\textsf{NP}} \)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow /> <mi mathvariant="sans-serif">NP</mi> </mmultiscripts> </math></EquationSource> </InlineEquation> in non-Pareto cases, respectively.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Handling concept drifts with traditional process discovery algorithms

  • Vincenzo Pasquadibisceglie

摘要

Event logs are data sets recording the executions (called cases) of a business process. Several process discovery algorithms have been defined to mine event logs and discover models of how activities of logged processes are being executed (activity traces). In the process discovery problems, the Pareto principle plays an important role. In fact, it is quite common that a large portion of log traces is held by a small fraction of top-frequent variants. Hence, accounting for the expected Pareto distribution of event logs, traditional process discovery algorithms are commonly used to discover process models by analyzing the prevalent trace behaviors. However, the Pareto principle is not always verified, especially in complex processes, where the majority of traces in the event log is often spanned on a high number of top-frequent trace-variants. In addition, traditional process discovery algorithms perform an offline analysis of event logs, assuming that logged processes remain in a steady state over time. But, the steady state is rarely the real-world case due to conceptual drifts. In this study, we use the traditional process discovery algorithms under the dynamic conditions of real-world processes. To this aim, we define two approaches, namely FAIRY \( ^{\textsf{P}} \) P and FAIRY \( ^{\textsf{NP}} \) NP , which detect drifts in the conformance of traces to process models and discover new process models on drifted traces. In FAIRY \( ^{\textsf{P}} \) P , a new process model is discovered on an extraction-based representation of a drift. In FAIRY \( ^{\textsf{NP}} \) NP , a new process model is discovered on an abstraction-based representation of a drift. The experimental results analyze the performance of the proposed approaches, also compared to a few related methods, showing the effectiveness of FAIRY \( ^{\textsf{P}} \) P in Pareto cases and FAIRY \( ^{\textsf{NP}} \) NP in non-Pareto cases, respectively.