A Clustering-Based Optimization Approach for Hospital Miscoding Correction
Bibliographic record
Abstract
This paper addresses the problem of correcting medical coding errors with respect to some coding recommendations. The problem consists in clustering medical codings and determining for each cluster the set of features to correct in order to maximize the financial benefits subject to coding correction effort constraints. For this purpose, we model the coding recommendation as a disjunction of hypercubes and introduce the concept of correction sets. A mixed integer linear programming model is then proposed to assign medical codes to correction sets in order to maximize the financial benefits. The miscoding is then explained by characterizing optimal clusters with association rules and coding error distribution. A case study on patient stays associated with malnutrition-related ICD codes is presented, and the performance of the proposed methodology is assessed in regard to the current coding staff practice. A significant increase in health services reimbursement is achieved with a limited number of subjects’ features reviewed. Note to Practitioners—Medical miscoding has a significant negative impact on hospitals with a financial loss for under coding and a penalty for over coding. Whether a medical review is necessary for all descriptive features of a miscoded subject? Is it possible to reduce unnecessary medical reviews without compromising the goal of increasing hospital financial benefits? This article attempts to answer these questions with a data-driven optimization approach to determine a limited number of miscoding clusters and the set of features to review for each in order to best balance the financial benefits and the medical review workload. The application to a real-life case study leads to a significant increase in hospital fiscal revenue of nearly 6,992,489.69 €, while reviewing only a small number of descriptive features (5293 out of 22056 features, or 24% of features). Causes are also provided for each discovered coding error subtype to ameliorate medical coders’ coding practices. Furthermore, the proposed approach allows the decision-maker to balance the cost-benefit and the requirement of public health institutions (i.e., miscoding rate).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".