Qualitative evaluation of clinician interaction with a machine learning algorithm for the assessment of patients with suspected acute heart failure in the emergency department
Bibliographic record
Abstract
Abstract Background N-terminal pro-B-type natriuretic peptide (NT-proBNP) assays have not been consistently implemented in practice despite being recommended in clinical guidelines for the assessment of acute heart failure. CODE-HF is a clinical decision-support tool that applies machine learning and NT-proBNP as a continuous measure and selected simple objective clinical variables to improve the diagnostic performance of NT-proBNP for acute heart failure. Purpose In a qualitative study, we aimed to explore the acceptance, and barriers and facilitators that led to positive clinician engagement with CODE-HF when used to assess anonymised clinical cases. Methods Individual semi-structured interviews were conducted either face-to-face or by video call with 17 clinicians from different disciplines working in Emergency Departments at 3 hospitals. They were asked to review five anonymised clinical cases and ‘think aloud’ about how they would assess the patient, and their interpretation of the CODE-HF metrics. These include a score of 0-100 representing an individualised probability of acute heart failure, diagnostic metrics and a classification of low, intermediate or high probability of acute heart failure (Figure 1). Interviews were audio recorded, transcribed and coded. Codes were mapped onto the four domains of the Unified Theory of Acceptance and Use of Technology model (performance expectancy, effort expectancy, social influences, facilitating conditions). Results Performance expectancy: Assessment could be improved using CODE-HF by facilitating objective communication between colleagues in a similar away to other widely used tools. The classification by probability score helped to reprioritise acute heart failure in cases where a diagnosis may have been missed. Effort expectancy: Statements relating to the positive or negative predictive value of a diagnosis of acute heart failure were viewed as useful information along with a visual traffic light system for the low-, intermediate- or high-probability categories. The absolute score was considered less useful to clinicians due to increased effort required for interpretation. Social influences: local and national guidelines carried the greatest weight of whether clinical decision support tools are used in practice, though respected research active colleagues and review on professional podcasts were also influential. Facilitating conditions: Access to a computer and clinical sample processing time were the only potential organisational issues identified as barriers. Clinicians were unanimous that clinical decision support tools provide supplementary information rather than replace clinical assessment which is central to the decision making process. Conclusion Clinicians reported that CODE-HF was a useful tool in the assessment of patients with breathlessness in the Emergency Department and identified the diagnostic metrics that were most helpful in guiding clinical decisions.CODE-HF display
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.083 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.005 | 0.007 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".