Machine Learning and Deep Learning Techniques for Prediction and Diagnosis of Leptospirosis: Systematic Literature Review
Bibliographic record
Abstract
Background: Leptospirosis, a zoonotic disease caused by Leptospira bacteria, continues to pose significant public health risks, particularly in tropical and subtropical regions. Objective: This systematic review aimed to evaluate the application of machine learning (ML) and deep learning (DL) techniques in predicting and diagnosing leptospirosis, focusing on the most used algorithms, validation methods, data types, and performance metrics. Methods: Using Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, Checklist for Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies (CHARMS), and Prediction model Risk of Bias Assessment Tool (PROBAST) tools, we conducted a comprehensive review of studies applying ML and DL models for leptospirosis detection and prediction, examining algorithm performance, data sources, and validation approaches. Results: Out of a total of 374 articles screened, 17 studies were included in the qualitative synthesis, representing approximately 4.5% of the initial pool. The review identified frequent use of algorithms such as support vector machines, artificial neural networks, decision trees, and convolutional neural networks (CNNs). Among the included studies, 88% (15/17) used traditional ML methods, and 24% (4/17) used DL techniques. Several models demonstrated high predictive performance, with reported accuracy rates ranging from 80% to 98%, notably with the U-Net CNN achieving 98.02% accuracy. However, public datasets were underused, with only 35% (6/17) of studies incorporating publicly available data sources; the majority (65%, 11/17) relied primarily on private datasets from hospitals, clinical records, or regional surveillance systems. Conclusions: ML and DL techniques demonstrate potential for improving leptospirosis prediction and diagnosis, but future research should focus on using larger, more diverse datasets, adopting transfer learning strategies, and integrating advanced ensemble and validation techniques to strengthen model accuracy and generalization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".