An Enhanced Approach to Predict Heart Disease for Genx using Machine Learning Algorithms
Bibliographic record
Abstract
Heart disease remains among the leading causes in which individuals die across the globe, especially the middle aged. It is not easy to predict it in the early days since there is a great number of various factors, which determine it, including the state of the clinical conditions, lifestyle behaviors, and processes, a lot of which are related to the working of the body. To aid this task this study proposes a new method of utilizing machine learning that is especially aimed at individuals belonging to the Gen X generation i.e., people between the age group of 40 to 60 years as they are the people that are more likely to experience an issue connected to the heart.The study relied on the UCI Heart Disease data and picked 216 records of individuals within this age category.The analysis of these records was done using 14 significant features of health care. To refine the data to facilitate the quality of the models made, the data were well prepared by addressing any irregular values, normalising the data on a normal scale and converting any non-numberical data to an appropriate form before the training of the models. There were four machine learning techniques that were evaluated, these include Logistic Regression, K-Nearest Neighbors (KNN), Decision Tree, and finally, Randome Forest.The study resulted that the KNN method performed best with the testing accuracy rating of 95 which is superior to the other methods.The research contributes to an existing gap in studies because this study targets a certain group of people of certain age as opposed to the general population. It also provides a comparison of the effectiveness of some of the commonly used machine learning methods to predict heart disease.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".