Enhancing Predictive Power of Cluster-Boosted Regression With Text-Based Indexing
Bibliographic record
Abstract
Clustering prior to regression analysis improves the accuracy of prediction in clinical decision making. However, most previously described methods focused on numerical data only. This paper investigated how well textual features can improve the accuracy of regression predictions. Preliminary diagnosis, diagnosis summary, and drug names used in prescriptions as provided in the MIMIC II dataset were used to derive textual features. We proposed the bag-of-entities indexing method, which relies on named entity recognition, a machine learning technique used for locating and identifying words into predefined classes. The proposed technique captured meaningful phrases from texts in health records and represented them in numerical vector format. Dimensionality of the data space was reduced using principal component analysis. The additional well-tuned textual features were then combined with existing numerical features in using cluster-boosted regression to predict patient mortality in ICU. The experimental results showed prediction improvement obtained from textual features over the use of numerical features only. We found that using the proposed indexing method outperformed traditional word-vector representation approaches (bag-of-words and bag-of-bigrams) as well as a state-of-the-art approach (Doc2vec) in terms of resulting accuracy in predicting death status. Moreover, instead of directly interpreting, the identifiable individual features were grouped into types and summarized. The summarized de-identified data of textual features handled by the proposed framework can support predictive classification while also reducing privacy concerns. Grouping of similar patients based on their electronic health records also benefits physicians through the improved differential diagnosis and effective treatment planning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".