A GeneSwarm-Enhanced Hybrid Ensemble Model for Predicting Cardiac Outcomes in Kawasaki Disease
Bibliographic record
Abstract
Kawasaki Disease (KD) is a paediatric vasculitis that may cause coronary artery lesions (CAL) in case it is not diagnosed and treated early. The problem of complicated clinical data and nonspecific markers has complicated proper forecasting of the long-term KD outcomes. This paper suggests a hybrid ensemble machine learning model to use XGBoost, AdaBoost, and Support Vector Machine (SVM) classifiers and soft voting in order to have robust and interpretable predictions. The GeneSwarm Feature Selector was used to select the features based on the Genetic Algorithms and Particle Swarm Optimization which can optimally identify the most informative and non-redundant clinical and laboratory features. The Babysitting algorithm was used to optimize hyperparameters dynamically and Interquartile Range (IQR) to eliminate outliers and Synthetic Minority Oversampling Technique (SMOTE) to deal with class imbalance were used during preprocessing. Assessed on 1,850 paediatric KD cases, the proposed model demonstrated a 97.10% accuracy, 96.50% precision, 96.80% recall, 96.60% F1-score, and a macro-AUC of 97.40% and is better than the current ML methods. Best predictive variables were Complications, Clinical Outcomes, Treatment Approach, Echocardiography, Laboratory Tests, Fever Duration, and Symptoms. These findings indicate that the framework is a reliable, interpretable, and clinically applicable instrument of early KD prognosis, which helps to plan personalized treatment and stratify risks. Through the incorporation of superior feature selection, ensemble learning, and a solid preprocessing, the research will add to an empirical and pragmatic method of enhancing cardiac care in pediatrics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".