Assessing Usefulness of the Dashboard Instrument to Review Equity (DIRE) Checklist to Evaluate Equity in Public Health Dashboards: Reliability Study
Bibliographic record
Abstract
BACKGROUND: The COVID-19 pandemic was a critical time for public health and though dashboards remained a source of critical health information for decision makers, key gaps in equity-based decision support were revealed. The Dashboard Instrument to Review Equity (DIRE) Framework and Checklist tool was developed to be a practical tool for public health departments to use in evaluating equity-based decision support mechanisms in their dashboards. OBJECTIVE: The objective of this agreement and reliability study was to validate the DIRE Checklist tool as a practical and reliable instrument for data practitioners to use in evaluating dashboards. METHODS: This study was divided into five steps to conduct the necessary analysis for agreement and reliability. Step 1 completed the development of the DIRE Checklist tool in Qualtrics. Step 2 focused on the parameters required for the selection of the 26 U.S. state-based dashboards. Step 3 was the user testing and assessment process for each reviewer to apply the DIRE tool to each dashboard. Step 4 was the different assessment methods conducted to specifically calculate the comparative analysis, inter-rater agreement, Intraclass Correlation Coefficients (ICC), and the cosine similarity for the Qualtrics, Reviewer, and Categorical scores. And lastly was Step 5 to conduct any qualitative assessment required on the notes. RESULTS: A total of 26 dashboards were evaluated with the DIRE Checklist tool by 2 reviewers. The overall percentage comparison for the Qualtrics Score was 31.7% for Reviewer 1 and 41.8% for Reviewer 2, resulting in a percent agreement of 72.7%. Additionally, the Categorical Scores saw substantial to high agreement across most categories for the percent agreement within each Category. The ICC score saw varying levels of agreement across different categories, with good agreement in the Qualtrics score. CONCLUSIONS: The reliability and agreement result of the study confirmed strong performance of the DIRE Checklist tool. The scores calculated were evaluated consistently and reliably by both raters-demonstrating the DIRE Checklist tool's ability to robustly evaluate different dashboards across a number of different categories and parameters.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.217 | 0.403 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.007 | 0.004 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".