An Efficient and User-Friendly Hybrid Approach for Classifying and Predicting Cardiovascular Disease Risks
Bibliographic record
Abstract
For any disease, health symptoms indicate its status in terms of severity whether medium, low, or high and their impact on life survival. There are stages of the disease that need to be classified. Among all diseases, severity is particularly significant in the case of heart disease. While health symptoms often suggest the presence of a disease, accurately predicting severity and categorizing risks, especially for complex conditions like cardiovascular disease (CVD), remains crucial. There are numerous factors that indicate the risk for cardiovascular problems. These include not only traditional risk factors like age and family history, but also genomic data, coronary artery scores, and carotid intima-media thickness analyzed through imaging studies, along with insights from biochemical inflammatory markers. Among these, warning signs such as shortness of breath, discomfort, fatigue, irregular heartbeat, leg numbness, and chest pain can affect the risk of a heart attack. The risks can be classified as modifiable or non-modifiable, time-frame risks, and low or severe risks. To predict and classify the risk in CVD, the integration of machine learning models and rule development can help determine the classification of risks. The hybrid approach is designed to identify and classify the risks into appropriate categories. These risks may affect health based on their severity. The computational process involved must be measured against accuracy and efficiency. The proposed hybrid model is expected to demonstrate exceptional performance compared to existing approaches, and these measures will be represented in graphical aids for better understanding and to ensure quality. This study explores a hybrid approach leveraging both machine learning and rule-based systems to achieve the objectives of the proposal.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".