Development of a primary care screening algorithm for the early detection of patients at risk of primary antibody deficiency
Bibliographic record
Abstract
BACKGROUND: Primary antibody deficiencies (PAD) are characterized by a heterogeneous clinical presentation and low prevalence, contributing to a median diagnostic delay of 3-10 years. This increases the risk of morbidity and mortality from undiagnosed PAD, which may be prevented with adequate therapy. To reduce the diagnostic delay of PAD, we developed a screening algorithm using primary care electronic health record (EHR) data to identify patients at risk of PAD. This screening algorithm can be used as an aid to notify general practitioners when further laboratory evaluation of immunoglobulins should be considered, thereby facilitating a timely diagnosis of PAD. METHODS: Candidate components for the algorithm were based on a broad range of presenting signs and symptoms of PAD that are available in primary care EHRs. The decision on inclusion and weight of the components in the algorithm was based on the prevalence of these components among PAD patients and control groups, as well as clinical rationale. RESULTS: We analyzed the primary care EHRs of 30 PAD patients, 26 primary care immunodeficiency patients and 58,223 control patients. The median diagnostic delay of PAD patients was 9.5 years. Several candidate components showed a clear difference in prevalence between PAD patients and controls, most notably the mean number of antibiotic prescriptions in the 4 years prior to diagnosis (5.14 vs. 0.48). The final algorithm included antibiotic prescriptions, diagnostic codes for respiratory tract and other infections, gastro-intestinal complaints, auto-immune symptoms, malignancies and lymphoproliferative symptoms, as well as laboratory values and visits to the general practitioner. CONCLUSIONS: In this study, we developed a screening algorithm based on a broad range of presenting signs and symptoms of PAD, which is suitable to implement in primary care. It has the potential to considerably reduce diagnostic delay in PAD, and will be validated in a prospective study. Trial registration The consecutive prospective study is registered at clinicaltrials.gov under NCT05310604.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".