Improving proactive primary healthcare after gestational diabetes: Protocol for implementation and evaluation of a quality improvement collaborative program in general practice (GooD4Mum) and baseline practice characteristics (Preprint)
Bibliographic record
Abstract
Abstract Background Gestational diabetes mellitus (GDM) is increasingly common, with short- and long-term health risks. Building on the GooD4Mum pilot, which demonstrated quality improvements in general practice for care after GDM, this project implemented and evaluated a primary care Quality Improvement Collaborative (QIC) program to optimize identification, recall, screening, and referral of patients after GDM. Objective This study aimed to assess the effectiveness of QIC activities relative to usual practice for improving general practice clinicians’ provision of follow-up and screening of patients with a history of GDM to ultimately support the onset of type 2 diabetes. Methods A 21-month, prospective non–randomized controlled trial (at practice level) was conducted, matching intervention practices 1:1 with controls. The QIC intervention is compared with care-as-usual in general practice to review implementation, effectiveness, and economic outcomes. For 18 months, the intervention practices engaged with the GooD4Mum QIC program, including education, training, resources, and Plan-Do-Study-Act (PDSA) cycles to implement locally relevant improvement activities. A clinical decision support system aided in the identification, screening, and tracking of patients with a history of GDM. Control practices provided care as usual. Primary outcomes are practice-level proportions of women with recorded type 2 diabetes screening, modifiable cardiometabolic risk factors, and referral to a diabetes prevention program. Secondary outcomes examine changes in care processes, adherence to clinical standards, and fidelity of intervention delivery. Outcome analyses use clinical data derived from practices via automated extraction. Baseline comparisons use t tests or chi-square tests. Primary and secondary outcomes are analyzed by repeated-measures ANOVA and/or cluster-adjusted generalized estimating equations, accounting for practice-level clustering, and sensitivity analyses include a per-protocol approach. Implementation outcomes are assessed through longitudinal qualitative interviews with practice leads, support staff, and stakeholders, guided by the CFIR (Consolidated Framework for Implementation Research) and the RE-AIM (Reach, Effectiveness, Adoption, Implementation and Maintenance) framework. Cost consequence analysis uses clinical activity and practice-reported data. Results The program ran in 9 general practices between April 2024 and September 2025, with data collection until December 2025. Practice characteristics and protocol deviation are summarized. Detailed outcome, cost consequence, and implementation process evaluations will be reported elsewhere. Conclusions The protocol offers a novel quality improvement approach enhanced by clinical decision support and automated data extraction to optimize identification, screening, and provision of lifestyle advice toward reducing type 2 diabetes after GDM. The evaluation is designed to generate actionable insights and a scalable implementation toolkit to improve care after gestational diabetes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.112 | 0.151 |
| Meta-epidemiology (narrow) | 0.004 | 0.006 |
| Meta-epidemiology (broad) | 0.006 | 0.010 |
| Bibliometrics | 0.005 | 0.008 |
| Science and technology studies | 0.010 | 0.004 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.005 | 0.006 |
| Research integrity | 0.007 | 0.011 |
| Insufficient payload (model declined to judge) | 0.054 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".