SnapLog turns kitchen videos into event logs via ViT frame embeddings, K-Means segmentation and R(2+1)D few-shot classification, reaching 90% and 85% Top-3
Synopsis
The work presents SnapLog, a pipeline that encodes each video frame with a pre-trained ViT, performs temporal segmentation through a frame-wise similarity matrix with K-Means and a greedy merge, then applies generalized few-shot classification with a 34-layer R(2+1)D encoder pre-trained on Sports-1M plus a linear head to convert raw video into timestamped event logs (either deterministic or uncertainty-aware logs that retain label probability distributions); on 16 brownie-baking videos from TUM Kitchen and 24 kitchen-cleaning videos from Epic Kitchens-100 it achieves frame-wise segmentation accuracy of 74.3% and 67.7%, and with augmentation plus 100 overlapping clips Top-3 classification accuracy of 90.1% and 85.
Fig. 1. SnapLog extracts logs from video data. Our method, SnapLog, constructs an event log from video data which can then be used downstream by standard process mining algorithms.
· Page 2Interpretation
It proposes a two-stage pipeline that turns video into event logs usable for process mining: ViT frame embeddings plus a similarity-matrix K-Means clustering yield atomic events, and a greedy merge folds too-short events into adjacent ones to form composite events. Compared with prior video-based process mining that relies on a fully known activity label dictionary and focuses on object tracking, or with VLM/LLM cascades requiring cloud compute, SnapLog targets a partially known label dictionary and resource-constrained on-premise deployment, and offers frame-level interpretable segmentation. On TUM Kitchen and Epic Kitchens-100, silhouette scores across k in {3,5,7} are compared, with k=7 giving 0.498 and 0.487, and frame-wise segmentation accuracy of 74.3% and 67.7%.
It assigns activity labels to the segmented clips with generalized few-shot classification, so only a few labeled samples are needed to adapt to a new environment. It brings few-shot video classification into video-to-log extraction, so adding new activity classes reduces to training a small linear head rather than requiring large labeled datasets or retraining the backbone. TUM and Epic Kitchens average only 11 and 7 segments per class; across four random seeds the standard deviation is only 2%-3%; moving from 10 non-overlapping to 100 overlapping clips improves Epic Kitchens Top-3 by up to 23%.
The pipeline can output uncertainty-aware event logs that retain the label probability distribution, not only a maximum-likelihood deterministic log. It explicitly preserves the label ambiguity inherent in video extraction and passes it downstream, connecting to recent work on uncertain and probabilistic event data instead of forcing a hard decision. With augmentation and 100 overlapping clips, Top-3 accuracy is 90.1% for TUM and 85.4% for Epic Kitchens, indicating the true activity lies within a narrow candidate set; the uncertain log can retain the three most likely labels and normalize them into a discrete probability distribution.
A largely sensible process model can be discovered from the extracted log, demonstrating end-to-end viability. Using a directly-follows graph obtained with Process Explorer from the TUM log, it shows video-derived logs can feed conventional process discovery, while a full quantitative evaluation of process discovery is stated by the authors to be out of scope. The authors report the model shows a generally correct brownie-baking procedure with sensible connections, alongside ordering inconsistencies such as Add Water followed by Get Water and Read Brownie Box appearing near the end.
Perspective
The work targets organizations that want to bring video recordings into process mining, and applies to settings where activities occur sequentially in time and each segment contains a single activity, such as manufacturing lines or kitchen operations; the method is designed for on-premise deployment on standard non-HPC hardware, needs roughly 15-20 videos of labeling effort to adapt to a new environment, and provides both deterministic logs and uncertainty-aware logs that retain label probabilities for downstream uncertainty-aware process analysis.
Segmentation and classification results come from two kitchen datasets, so behavior in other domains and more complex scenes still needs testing; the authors note the current assumption that only one activity appears per segment, that only activity labels and temporal information are extracted while spatial information is lost, and that the modular design is not yet end-to-end differentiable; a quantitative evaluation of process discovery is stated to be out of scope, so fidelity from log to process model remains an open question; initial annotation is still manual, with VLMs in the loop suggested as future work.
