Evaluation Methods, Indicators, and Outcomes in Learning Health Systems: Protocol for a Jurisdictional Scan
Bibliographic record
Abstract
BACKGROUND: In learning health systems (LHSs), real-time evidence, informatics, patient-provider partnerships and experiences, and organizational culture are combined to conduct "learning cycles" that support improvements in care. Although the concept of LHSs is fairly well established in the literature, evaluation methods, mechanisms, and indicators are less consistently described. Furthermore, LHSs often use "usual care" or "status quo" as a benchmark for comparing new approaches to care, but disentangling usual care from multifarious care modalities found across settings is challenging. There is a need to identify which evaluation methods are used within LHSs, describe how LHS growth and maturity are conceptualized, and determine what tools and measures are being used to evaluate LHSs at the system level. OBJECTIVE: This study aimed to (1) identify international examples of LHSs and describe their evaluation approaches, frameworks, indicators, and outcomes; and (2) describe common characteristics, emphases, assumptions, or challenges in establishing counterfactuals in LHSs. METHODS: A jurisdictional scan, which is a method used to explore, understand, and assess how problems have been framed by others in a given field, will be conducted according to modified PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. LHSs will be identified through a search of peer-reviewed and gray literature using Ovid MEDLINE, EBSCO CINAHL, Ovid Embase, Clarivate Web of Science, PubMed non-MEDLINE databases, and the web. We will describe evaluation approaches used both at the LHS learning cycle and system levels. To gain a comprehensive understanding of each LHS, including details specific to evaluation, self-identified LHSs will be included if they are described according to at least 4 of 11 prespecified criteria (core functionalities, analytics, use of evidence, co-design or implementation, evaluation, change management or governance structures, data sharing, knowledge sharing, training or capacity building, equity, and sustainability). Search results will be screened, extracted, and analyzed to inform a descriptive review pertaining to our main objectives. Evaluation methods and approaches, both within learning cycles and at the system level, as well as frameworks, indicators, and target outcomes, will be identified and summarized descriptively. Across evaluations, common challenges, assumptions, contextual factors, and mechanisms will be described. RESULTS: As of October 2024, the database searches described above yielded 3503 citations after duplicate removal. Full-text screening of 117 articles is complete, and 49 articles are under analysis. Results are expected in early 2025. CONCLUSIONS: This research will characterize the current landscape of LHS evaluation approaches and provide a foundation for developing consistent and scalable metrics of LHS growth, maturity, and success. This work will also serve to identify opportunities for improving the alignment of current evaluation approaches and metrics with population health needs, community priorities, equity, and health system strategic aims. TRIAL REGISTRATION: Open Science Framework b5u7e; https://osf.io/b5u7e. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/57929.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.204 | 0.262 |
| Meta-epidemiology (narrow) | 0.005 | 0.005 |
| Meta-epidemiology (broad) | 0.011 | 0.012 |
| Bibliometrics | 0.013 | 0.020 |
| Science and technology studies | 0.007 | 0.006 |
| Scholarly communication | 0.009 | 0.009 |
| Open science | 0.006 | 0.007 |
| Research integrity | 0.011 | 0.010 |
| Insufficient payload (model declined to judge) | 0.105 | 0.018 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".