Machine Learning Assisted MALDI Mass Spectrometry for Rapid Antimicrobial Resistance Prediction in Clinicals
Bibliographic record
Abstract
Antimicrobial susceptibility testing (AST) plays a critical role in assessing the resistance of individual microbial isolates and determining appropriate antimicrobial therapeutics in a timely manner. However, conventional AST normally takes up to 72 h for obtaining the results. In healthcare facilities, the global distribution of vancomycin-resistant Enterococcus fecium (VRE) infections underscores the importance of rapidly determining VRE isolates. Here, we developed an integrated antimicrobial resistance (AMR) screening strategy by combining matrix-assisted laser desorption ionization mass spectrometry (MALDI-MS) with machine learning to rapidly predict VRE from clinical samples. Over 400 VRE and vancomycin-susceptible E. faecium (VSE) isolates were analyzed using MALDI-MS at different culture times, and a comprehensive dataset comprising 2388 mass spectra was generated. Algorithms including the support vector machine (SVM), SVM with L1-norm, logistic regression, and multilayer perceptron (MLP) were utilized to train the classification model. Validation on a panel of clinical samples (external patients) resulted in a prediction accuracy of 78.07%, 80.26%, 78.95%, and 80.54% for each algorithm, respectively, all with an AUROC above 0.80. Furthermore, a total of 33 mass regions were recognized as influential features and elucidated, contributing to the differences between VRE and VSE through the Shapley value and accuracy, while tandem mass spectrometry was employed to identify the specific peaks among them. Certain ribosomal proteins, such as A0A133N352 and R2Q455, were tentatively identified. Overall, the integration of machine learning with MALDI-MS has enabled the rapid determination of bacterial antibiotic resistance, greatly expediting the usage of appropriate antibiotics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".