Mining infrequent group of motifs from multidimensional time series: A case study at Alfa Laval AB
Bibliographic record
Abstract
In collaboration with an industrial partner, Alfa Laval AB, this thesis discusses a novel approach for identifying operational modes, specifically a cleaning mode, in separator machines without the benefit of labelled data and with very limited operating knowledge. Understanding the operational modes is crucial for comprehending the machine’s behaviour and ensuring its optimal performance. Alfa Laval AB relies on a threshold-based fault detection system. The cleaning mode triggers vibrations that confuse the machine’s fault detection system, resulting in false alarms. The primary challenge revolves around the limited understanding of this infrequent cleaning mode, occurring periodically for 1-2 hours at intermittent intervals. To tackle this, we approach the problem as a data mining task. Matrix Profile (MP), a powerful tool in time series data analysis excels at identifying motifs and discords but struggles to distinguish between frequent and non-frequent motifs. To address the drawback, we introduced an innovative approach capable of detecting frequent motifs and non-frequent motifs from the matrix profile output. The fundamental concept involves extracting the top-K motif matches using the Matrix Profile (MP) and systematically monitoring the evolution of structural similarity through pairwise similarity matrix calculation, progressing from pairs of two motifs to a group of K motifs. This approach helps us to identify infrequent motifs that contain the most similar patterns which will be a good fit to address our challenge of identifying the cleaning mode.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".