Lifestyle and occupational risk assessment using machine learning methods and development a web-based nomogram for prediction of bladder cancer Short title: Risk assessment and nomogram to predict bladder cancer
Bibliographic record
Abstract
Abstract Background: Urinary bladder cancer (UBC) is one of the most disabling yet prevalent types of cancer. With its burden on the rise, UBC is becoming more prevalent each year and the numbers will not stop until we decide to make an endeavor to increase our knowledge even further and to reduce the risk factors that predispose to it. We tried to use machine learning (ML) methods to develop a risk prediction model in this study.Methods: This population-based case-control study is focused on 692 cases of bladder cancer, and 692 healthy people. The ML including Neural Network (NN), Random Forest (RF), Regression Tree (RT), Naive Bayes (NB), and Logistic Regression (LR) were applied. Area Under Curve (AUC) of receiver operating characteristic, precision, sensitivity, specificity, F1 and recall were used to evaluate the model performance. Finally, a dynamic web-based nomogram was constructed to calculate the probability of bladder cancer.Results: The RF (AUC=.98, precision=91.3%) and NN (AUC=.97, precision=91.5%) had the best performance and the RT (AUC=0.98, precision=89.5%) was in the next rank. Based on variable importance analysis in RF, that recurrence infection, smoking, neurogenic bladder, low diet on fruit and vegetable, can and ham using were respectively the most significant variables which effect on probability of bladder cancer.Conclusion: Based on ML findings, urinary bladder risk factors can be classified into two major groups: exogenous factors such as smoking, dietary factors, environmental risks, occupational risks, and previous medical conditions of a patient. The second group is the endogenous or intrinsic factors that include gender, age, and genetics of a person. The presented web-based nomogram can be considered as a user-friendly clinical tool to predict the probability of UBC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".