Application of machine learning for predicting the incubation period of water droplet erosion in metals
Bibliographic record
Abstract
Water droplet erosion (WDE) is a critical degradation phenomenon that significantly affects component lifespan and performance in power generation, aerospace, and wind energy industries. The incubation period—the initial phase before visible material loss occurs—is particularly crucial for maintenance planning and material selection yet remains challenging to predict accurately due to the complex interplay of material properties and impact conditions. Traditional empirical models have shown limited predictive capability due to their reliance on numerous adjustable parameters with insufficient physical interpretation. This study aimed to develop and validate a machine learning (ML) approach for accurately predicting the WDE incubation period across different metallic materials and impact conditions. The performance of various ML algorithms is evaluated while investigating the effect of data transformation techniques on prediction accuracy. A range of ML models—linear regression (LR), decision tree regressor (DT), random forest regressor (RF), gradient boosting regressor (GBR), and artificial neural networks (ANN)—were trained and validated using experimental data from five different alloys under various impact conditions. Data transformation methods significantly enhanced model performance, with the LR model using Box-Cox transformation achieving the highest accuracy (R2 > 90%, low MAE), followed by the ANN model with Yeo-Johnson transformation (R2 > 85%). Feature importance analysis through SHAP values revealed that impact velocity and surface hardness were the most influential factors affecting incubation period, providing valuable physical insights into the erosion mechanism. Hyperparameter optimization techniques showed minimal improvement in model performance, suggesting that the transformations effectively captured the underlying relationships in the data. This research represents the first comprehensive application of ML techniques to WDE incubation period prediction, establishing a methodological framework that integrates experimental data, statistical analysis, and advanced ML algorithms. Unlike previous approaches, our methodology (1) systematically evaluates multiple ML algorithms and transformation techniques for WDE prediction, (2) provides quantitative assessment of feature importance that aligns with physical understanding of erosion mechanisms, (3) demonstrates superior predictive accuracy compared to traditional empirical models, and (4) offers a generalizable approach applicable across different metallic materials and impact conditions. This work bridges the gap between data-driven modeling and physical understanding of WDE, providing a valuable tool for engineers to optimize material selection and maintenance strategies in erosion-prone applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".