Fault Prediction in Energy Systems: Telemetric Data and Machine Learning Approaches
Bibliographic record
Abstract
The uninterrupted and reliable provision of energy is of paramount importance for the sustainability of economic activities and the welfare of society. The intricate structures formed by numerous components within the energy infrastructure create a basis for failures to yield severe consequences. This study aims to develop predictive maintenance models using machine learning algorithms by examining telemetry data, such as voltage, rotational speed, pressure, and vibration obtained from machinery, in conjunction with historical failure and maintenance records. The dataset utilized in this study comprises hourly telemetry information, including voltage, revolutions per minute, pressure, and vibration, collected from 100 distinct machines throughout the year 2015. Furthermore, supplementary information, such as the machines past failure occurrences, maintenance logs, and technical specifications, has been considered in the development of the models. In this study, timedependent patterns derived from sensor data, along with historical maintenance and failure information, have been integrated and analyzed using machine learning algorithms, namely Stochastic Gradient Descent Classifier (SGDClassifier), eXtreme Gradient Boosting (XGBoost), and Histogram-based Gradient Boosting (HGBClassifier). The results have been analyzed based on performance metrics such as Precision, Recall and F1-Score. Notably, the XGBoost and HGBClassifier algorithms demonstrated superior performance in early failure detection, particularly within 24-and 48 -hour prediction windows, achieving high F1-Score values. Consequently, this approach aims to transcend the traditional reactive maintenance paradigm, thereby enhancing operational efficiency and preventing unforeseen downtimes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".