A study of the enablers and barriers to the collection of sociodemographic data by public health units in Ontario, Canada during the COVID-19 pandemic
Bibliographic record
Abstract
BACKGROUND: Collection and use of sociodemographic data (SDD), including race, ethnicity and income, are foundational to understanding health inequities. Ontario's public health units collected SDD as part of COVID-19 case management and vaccination activities. This research aimed to identify enablers and barriers to collecting SDD during COVID-19 case management and vaccination. METHODS: As part of a larger mixed-method research study [1], qualitative methods were used to identify enablers and barriers to SDD collection during the COVID-19 pandemic. Purposive sampling was used to recruit participants from Ontario's 34 public health units. Sixteen focus groups and eight interviews were conducted virtually using Zoom. Interview data were transcribed and analyzed using inductive and deductive qualitative description. RESULTS: SDD collection enablers included: legally mandating SDD collection and having dedicated data systems, technological and legal supports, senior management championing SDD collection, establishing rapport and trust between staff and clients, and gaining insight from the experiences from local communities and other jurisdictions. Identified barriers to SDD collection included: provincial data systems being perceived as lacking user-friendliness, SDD collection "was not a priority," time and other constraints on building staff and client rapport, and perceived discomfort with asking and answering personal SDD questions. CONCLUSION: A combination of provincial and local organizational strategies including supportive data systems, training, and frameworks for data collection and use, are needed to normalize and scale up SDD collection by local health units beyond the context of the COVID-19 pandemic.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.019 | 0.006 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".