Optimizing Performance of Equipment Fleets Under Dynamic Operating Conditions: Generalizable Shift Detection and Multimodal LLM-Assisted State Labeling
Bibliographic record
Abstract
This paper presents OpS-EWMA-LLM (Operational State Shifts Detection using Exponential Weighted Moving Average and Labeling using Large Language Model), a hybrid framework that combines fleet-normalized statistical shift detection with LLM-assisted diagnostics to identify and interpret operational state changes across heterogeneous fleets. First, we introduce a residual-based EWMA control chart methodology that uses deviations of each component’s sensor reading from its fleet-wide expected value to detect anomalies. This statistical approach yields near-zero false negatives and flags incipient faults earlier than conventional methods, without requiring component-specific tuning. Second, we implement a pipeline that integrates an LLM with retrieval-augmented generation (RAG) architecture. Through a three-phase prompting strategy, the LLM ingests time-series anomalies, domain knowledge, and contextual information to generate human-interpretable diagnostic insights. Finaly, unlike existing approaches that treat anomaly detection and diagnosis as separate steps, we assign to each detected event a criticality label based on both statistical score of the anomaly and semantic score from the LLM analysis. These labels are stored in the OpS-Vector to extend the knowledge base of cases for future retrieval. We demonstrate the framework on SCADA data from a fleet of wind turbines: OpS-EWMA successfully identifies critical temperature deviations in various components that standard alarms missed, and the LLM (augmented with relevant documents) provides rationalized explanations for each anomaly. The framework demonstrated robust performance and outperformed baseline methods in a realistic zero-tuning deployment across thousands of heterogeneous equipment units operating under diverse conditions, without component-specific calibration. By fusing lightweight statistical process control with generative AI, the proposed solution offers a scalable, interpretable tool for condition monitoring and asset management in Industry 4.0/5.0 settings. Beyond its technical contributions, the outcome of this research is aligned with the UN Sustainable Development Goals SDG 7, SDG 9, SDG 12, SDG 13.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".