MétaCan
Menu
Back to cohort
Record W4321770315 · doi:10.1109/tase.2023.3247177

A Clustering-Based Optimization Approach for Hospital Miscoding Correction

2023· article· en· W4321770315 on OpenAlexaff
Chen He, Benjamin Dalmas, Cédric Bousquet, Béatrice Trombert‐Paviot, Xiaolan Xie

Bibliographic record

VenueIEEE Transactions on Automation Science and Engineering · 2023
Typearticle
Languageen
FieldHealth Professions
TopicMedical Coding and Health Information
Canadian institutionsComputer Research Institute of Montréal
FundersNational Natural Science Foundation of China
KeywordsMedical classificationComputer scienceCluster analysisCoding (social sciences)Linear programmingData miningArtificial intelligenceAlgorithmMedicineStatisticsMathematics

Abstract

fetched live from OpenAlex

This paper addresses the problem of correcting medical coding errors with respect to some coding recommendations. The problem consists in clustering medical codings and determining for each cluster the set of features to correct in order to maximize the financial benefits subject to coding correction effort constraints. For this purpose, we model the coding recommendation as a disjunction of hypercubes and introduce the concept of correction sets. A mixed integer linear programming model is then proposed to assign medical codes to correction sets in order to maximize the financial benefits. The miscoding is then explained by characterizing optimal clusters with association rules and coding error distribution. A case study on patient stays associated with malnutrition-related ICD codes is presented, and the performance of the proposed methodology is assessed in regard to the current coding staff practice. A significant increase in health services reimbursement is achieved with a limited number of subjects’ features reviewed. Note to Practitioners—Medical miscoding has a significant negative impact on hospitals with a financial loss for under coding and a penalty for over coding. Whether a medical review is necessary for all descriptive features of a miscoded subject? Is it possible to reduce unnecessary medical reviews without compromising the goal of increasing hospital financial benefits? This article attempts to answer these questions with a data-driven optimization approach to determine a limited number of miscoding clusters and the set of features to review for each in order to best balance the financial benefits and the medical review workload. The application to a real-life case study leads to a significant increase in hospital fiscal revenue of nearly 6,992,489.69 €, while reviewing only a small number of descriptive features (5293 out of 22056 features, or 24% of features). Causes are also provided for each discovered coding error subtype to ameliorate medical coders’ coding practices. Furthermore, the proposed approach allows the decision-maker to balance the cost-benefit and the requirement of public health institutions (i.e., miscoding rate).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.961
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.099
GPT teacher head0.373
Teacher spread0.274 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueIEEE Transactions on Automation Science and EngineeringSame topicMedical Coding and Health InformationFrench-language works237,207