Machine learning approach to the assessment and prediction of solid particle erosion of metals
Bibliographic record
Abstract
Solid particle erosion (SPE) is a tribological phenomenon in which a surface is impacted by a stream of particles, causing gradual removal of material. This poses significant challenges in aerospace, particularly when operating in harsh environments. Despite decades of data gathering and empirical model development, accurately predicting SPE remains challenging due to the complexity of the phenomenon and the variability in testing conditions. In this study, we compiled a database of over 1,000 erosion tests on metals from existing studies and internal experiments, noting material properties, test conditions, and literature metadata. Machine learning (ML) models, including Random Forest, Neural Networks, Support Vector Regression, and XGBoost were employed to predict erosion rates. XGBoost was most performant, achieving a mean absolute error of 15-16% on test data. Model performance was further validated by predicting results published in the ASTM G76 standard; predictions were within the interlaboratory standard deviation for tests at 70 m/s. Feature importance and partial dependence plots were used to evaluate the influence of different variables on erosion predictions. While particle velocity, particle size, and impact angle show the expected influence, features such as target density and Poisson’s ratio showed exaggerated effects due to their role in classifying outlier materials. These results show the promise of ML for SPE prediction across a range of conditions and suggest that the broader erosion literature is valuable for quantitative predictions, while also acknowledging limitations in the ML approach, particularly where data sparsity and feature correlations hinder the accurate assessment of feature influence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".