MétaCan
Menu
Back to cohort
Record W4415728343 · doi:10.2196/71094

Assessing Usefulness of the Dashboard Instrument to Review Equity (DIRE) Checklist to Evaluate Equity in Public Health Dashboards: Reliability Study

2025· article· en· W4415728343 on OpenAlexvenueno aff
Paulina Sosa, Emir A. Syailendra, Harold P. Lehmann, Hadi Kharrazi

Bibliographic record

VenueJMIR Public Health and Surveillance · 2025
Typearticle
Languageen
FieldHealth Professions
TopicPublic Health Policies and Education
Canadian institutionsnot available
Fundersnot available
KeywordsChecklistEquity (law)Public healthReliability (semiconductor)DashboardData collection

Abstract

fetched live from OpenAlex

BACKGROUND: The COVID-19 pandemic was a critical time for public health and though dashboards remained a source of critical health information for decision makers, key gaps in equity-based decision support were revealed. The Dashboard Instrument to Review Equity (DIRE) Framework and Checklist tool was developed to be a practical tool for public health departments to use in evaluating equity-based decision support mechanisms in their dashboards. OBJECTIVE: The objective of this agreement and reliability study was to validate the DIRE Checklist tool as a practical and reliable instrument for data practitioners to use in evaluating dashboards. METHODS: This study was divided into five steps to conduct the necessary analysis for agreement and reliability. Step 1 completed the development of the DIRE Checklist tool in Qualtrics. Step 2 focused on the parameters required for the selection of the 26 U.S. state-based dashboards. Step 3 was the user testing and assessment process for each reviewer to apply the DIRE tool to each dashboard. Step 4 was the different assessment methods conducted to specifically calculate the comparative analysis, inter-rater agreement, Intraclass Correlation Coefficients (ICC), and the cosine similarity for the Qualtrics, Reviewer, and Categorical scores. And lastly was Step 5 to conduct any qualitative assessment required on the notes. RESULTS: A total of 26 dashboards were evaluated with the DIRE Checklist tool by 2 reviewers. The overall percentage comparison for the Qualtrics Score was 31.7% for Reviewer 1 and 41.8% for Reviewer 2, resulting in a percent agreement of 72.7%. Additionally, the Categorical Scores saw substantial to high agreement across most categories for the percent agreement within each Category. The ICC score saw varying levels of agreement across different categories, with good agreement in the Qualtrics score. CONCLUSIONS: The reliability and agreement result of the study confirmed strong performance of the DIRE Checklist tool. The scores calculated were evaluated consistently and reliably by both raters-demonstrating the DIRE Checklist tool's ability to robustly evaluate different dashboards across a number of different categories and parameters.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.217
metaresearch head score (Gemma)0.403
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.783
Threshold uncertainty score0.966

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2170.403
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.003
Bibliometrics0.0070.004
Science and technology studies0.0020.002
Scholarly communication0.0030.003
Open science0.0020.005
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.237
GPT teacher head0.546
Teacher spread0.309 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designObservational
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Public Health and SurveillanceSame topicPublic Health Policies and EducationFrench-language works237,207