Utilizing artificial intelligence for the detection of hemarthrosis in hemophilia using point-of-care ultrasonography
Bibliographic record
Abstract
Background: Recurrent hemarthrosis and resultant hemophilic arthropathy are significant causes of morbidity in persons with hemophilia, despite the marked evolution of hemophilia care. Prevention, timely diagnosis, and treatment of bleeding episodes are key. However, a physical examination or a patient's assessment of musculoskeletal pain may not accurately identify a joint bleed. This difficulty is compounded as hemophilic arthropathy progresses. Objectives: Our system aims to utilize artificial intelligence and ultrasonography (US; point-of-care and handheld) to enable providers, and ultimately patients, to detect joint bleeds at the bedside and at home. We aimed to develop and assess the reliability of artificial intelligence algorithms in detecting and segmenting synovial recess distension (SRD; an indicator of disease activity) on US images of adult and pediatric knee, elbow, and ankle joints. Methods: A total of 12,145 joint exams, comprising 61,501 US images from 7 international healthcare centers, were collected. The dataset included healthy participants and adult and pediatric persons with hemophilia, with and without SRD. Images were manually labeled by 2 experts and used to train binary convolutional neural network classifiers and segmentation models. Metrics to evaluate performance included accuracy, sensitivity, specificity, and area under the curve. Results: The algorithms exhibited high performance across all joints and all cohorts. Specifically, the knee model showed an accuracy of 97%, sensitivity of 96%, specificity of 97%, and an area under the curve of 0.97 in SRD. High Dice coefficients (80%-85%) were achieved in segmentation tasks across all joints. Conclusion: This technology could assist with the early detection and management of hemarthrosis in hemophilia.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".