Skip to main content
Back to timeline
arXivSource publication:

Across 14 event logs, remaining-time error and uncertainty rise with delay, and uncertainty features lift average delay-detection recall from 0.21 to 0.61

Synopsis

Analyzing the intrinsic difficulty of delay detection across 14 public event logs, the study finds that remaining times are strongly right-skewed, that existing LSTM models capture the mode of the distribution but perform poorly on high-delay cases, and that prediction error and prediction-interval width both grow with true remaining time (positive in 10 of 14 logs); it then evaluates imbalanced regression approaches (SMOGN resampling, CSW, BMSE, EAL, SERA) with only limited benefit, and instead feeds uncertainty features from a survival model (prediction standard deviation, 80% and 90% prediction-interval widths, tail mass, and temporal context) into a CatBoost classifier for binary delay detection, raising average recall from 0.21 to 0.61 at the q=0.

Source-provided article image: Mind the Long Tail: Understanding the Difficulty of Delay Detection in Business Processes

Interpretation

The work characterizes where delay detection is hard: remaining-time distributions are strongly right-skewed, models capture the mode but fail on high-delay cases, and predictive uncertainty grows with remaining time. Prior work typically evaluates remaining-time prediction with aggregate metrics such as MAE, which hide performance differences across the target distribution; this work decomposes error into many/medium/few regions and systematically reports heteroscedasticity. Analysis over 14 public event logs reports the Fisher moment coefficient of skewness (e.g., 11.5 for BPIC20DD, 10.6 for BPIC20RFP, 10.0 for BPIC15-3, and -0.1 for Helpdesk) and Spearman correlations between true remaining time and both absolute error and survival-model prediction-interval width (positive for both measures in 10 of 14 logs, reaching 0.64 and 0.80 on BPIC20ID, while BPIC15-2 and Helpdesk are negative).

Data-level and algorithm-level approaches to target imbalance yield limited benefit for long-tail prediction, suggesting tail data scarcity may not be the primary bottleneck. Imbalanced regression methods (SMOGN, CSW, BMSE, EAL, SERA) were previously validated mainly in classification, vision, or time-series settings; this work systematically compares them for PPM remaining-time regression and adapts SMOGN neighbor selection to left-padded prefix sequences. Across 14 logs, SMOGN reduces few-shot nMAE by 8.0% but increases medium and many-shot error by 20.5% and 31.4%, raising overall nMAE by 17.5%; among algorithm-level methods, SERA ranks best in the few-shot region (1.00) and significantly outperforms Vanilla, CSW, and EAL there, yet is significantly worse overall, while EAL and CSW achieve the best overall ranks; a Friedman test shows p<0.001 and pairwise comparisons use the Wilcoxon signed-rank test.

Using uncertainty estimates from a survival model as features for a downstream classifier substantially improves identification of high-delay cases. Prior uncertainty-aware remaining-time prediction work focused on quantifying uncertainty itself, without studying the interaction between skewed target distributions and uncertainty or using uncertainty for the decision task of delay detection. Across 14 logs, five random seeds, and a q=0.8 threshold, average recall rises from 0.21 to 0.61, average F1 from 0.21 to 0.44, and average PR-AUC from 0.23 to 0.33, while average precision moves from 0.38 to 0.42 with no significant overall difference (p=0.17); improvements in recall, F1, and PR-AUC correspond to p<0.0001, p<0.0001, and p<0.0003; on BPIC20PTC the baseline identifies only 1 of 175 high-delay cases while the uncertainty-aware model detects 132; similar improvements appear under a stricter q=0.9 definition.

The uncertainty-aware model detects delayed cases earlier and more reliably during case execution. The work normalizes prefix length by the total number of events in the full trace and evaluates performance by execution stage, characterizing the temporal availability of delay detection rather than only final classification performance. On prefix subsets corresponding to up to 20%, 40%, 60%, and 80% of process execution, the uncertainty-aware model outperforms the baseline at all prefix ratios, with the most pronounced gains at early stages: recall is substantially higher even when only a small portion of the case is observed, precision remains comparable, and F1 and PR-AUC improve consistently.

Perspective

The results target organizations that use event logs for predictive process monitoring, in settings where remaining-time regression underpins the need to identify high-delay cases in advance; methodologically the pipeline uses a data-aware LSTM encoder with a survival-analysis head, delay detection is framed as binary classification against a quantile threshold (e.g., q=0.8, with q=0.9 also examined), uncertainty features include prediction standard deviation, 80% and 90% prediction-interval widths, tail mass, and temporal context such as elapsed time, and a CatBoost classifier is used downstream. The authors note that delay detection abstracts the regression problem into threshold-based classification, so these gains should not be read as improvements in remaining-time estimation itself; next steps include richer uncertainty modeling, additional contextual information, and online adaptation of delay detection in changing process environments.

The direction and strength of heteroscedasticity are not uniform across processes: BPIC15-2 and Helpdesk show negative correlations for both measures, while BPIC17W and BPIC15-4 show mixed behavior, so the usefulness of uncertainty as a delay signal may depend on the specific process. The underlying sources of this uncertainty are not yet fully disentangled and may stem from process variability, missing context, or concept drift. Delay detection depends on a threshold choice and loses information, and the authors suggest modeling the full conditional distribution of remaining time so that delay detection and related problems such as service level agreement violations can be formulated as downstream decision rules. In addition, the evaluation is based on public event logs and LSTM-based models, so how strongly the findings generalize to other domains, industrial settings, and other architectures remains to be tested.

Sources