MétaCan
Menu
Back to cohort
Record W2589364158 · doi:10.1186/s13643-017-0431-9

Identification of validated case definitions for chronic disease using electronic medical records: a systematic review protocol

2017· review· en· W2589364158 on OpenAlexaff
Sepideh Souri, Nicola E. Symonds, Azin Rouhi, Brendan Cord Lethebe, Stephanie Garies, Paul E. Ronksley, Tyler Williamson, Gabriel E. Fabreau, Richard Birtwhistle, Hude Quan, Kerry McBrien

Bibliographic record

VenueSystematic Reviews · 2017
Typereview
Languageen
FieldMedicine
TopicChronic Disease Management Strategies
Canadian institutionsQueen's UniversityUniversity of British ColumbiaUniversity of AlbertaUniversity of Calgary
Fundersnot available
KeywordsMedicineProtocol (science)UsabilityIdentification (biology)Health careSet (abstract data type)Systematic reviewHealth informaticsData sciencePopulationMEDLINEData miningMedical emergencyComputer sciencePublic healthAlternative medicineNursingPathology

Abstract

fetched live from OpenAlex

BACKGROUND: Primary care electronic medical record (EMR) data are being used for research, surveillance, and clinical monitoring. To broaden the reach and usability of EMR data, case definitions must be specified to identify and characterize important chronic conditions. The purpose of this study is to identify all case definitions for a set of chronic conditions that have been tested and validated in primary care EMR and EMR-linked data. This work will provide a reference list of case definitions, together with their performance metrics, and will identify gaps where new case definitions are needed. METHODS: We will consider a set of 40 chronic conditions, previously identified as potentially important for surveillance in a review of multimorbidity measures. We will perform a systematic search of the published literature to identify studies that describe case definitions for clinical conditions in EMR data and report the performance of these definitions. We will stratify our search by studies that use EMR data alone and those that use EMR-linked data. We will compare the performance of different definitions for the same conditions and explore the influence of data source, jurisdiction, and patient population. DISCUSSION: EMR data from primary care providers can be compiled and used for benefit by the healthcare system. Not only does this work have the potential to further develop disease surveillance and health knowledge, EMR surveillance systems can provide rapid feedback to participating physicians regarding their patients. Existing case definitions will serve as a starting point for the development and validation of new case definitions and will enable better surveillance, research, and practice feedback based on detailed clinical EMR data. SYSTEMATIC REVIEW REGISTRATION: PROSPERO CRD42016040020.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.143
metaresearch head score (Gemma)0.176
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.857
Threshold uncertainty score0.757

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1430.176
Meta-epidemiology (narrow)0.0060.006
Meta-epidemiology (broad)0.0180.016
Bibliometrics0.0260.019
Science and technology studies0.0050.005
Scholarly communication0.0070.010
Open science0.0060.006
Research integrity0.0080.006
Insufficient payload (model declined to judge)0.0370.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.329
GPT teacher head0.511
Teacher spread0.182 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSystematic review
DomainMethods
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2017
Admission routes1
Has abstractyes

Explore more

Same venueSystematic ReviewsSame topicChronic Disease Management StrategiesFrench-language works237,207