Sociodemographic data collection for health equity measurement: a mixed methods study examining public opinions
Bibliographic record
Abstract
Monitoring inequalities in healthcare is increasingly being recognized as a key first step in providing equitable access to quality care. However, the detailed sociodemographic data that are necessary for monitoring are currently not routinely collected from patients in many jurisdictions. We undertook a mixed methods study to generate a more in-depth understanding of public opinion on the collection of patient sociodemographic information in healthcare settings for equity monitoring purposes in Ontario, Canada. The study included a provincial survey of 1,306 Ontarians, and in-depth interviews with a sample of 34 individuals. Forty percent of survey participants disagreed that it was important for information to be collected in healthcare settings for equity monitoring. While there was a high level of support for the collection of language, a relatively large proportion of survey participants felt uncomfortable disclosing household income (67%), sexual orientation (40%) and educational background (38%). Variation in perceived importance and comfort with the collection of various types of information was observed among different survey participant subgroups. Many in-depth interview participants were also unsure of the importance of the collection of sociodemographic information in healthcare settings and expressed concerns related to potential discrimination and misuse of this information. Study findings highlight that there is considerable concern regarding disclosure of such information in healthcare settings among Ontarians and a lack of awareness of its purpose that may impede future collection of such information. These issues point to the need for increased education for the public on the purpose of sociodemographic data collection as a strategy to address this problem, and the use of data collection strategies that reduce discomfort with disclosure in healthcare settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.084 | 0.068 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".