Sociodemographic Data Collection in Healthcare Settings
Bibliographic record
Abstract
BACKGROUND: Federal, provincial, and municipal organizations in Canada have recently begun to promote an equity agenda for their health systems, but much of the necessary data by which to identify those with social disadvantage are not currently collected. METHODS: We conducted a national survey of 1005 Canadian adults to assess the perceived importance of, and concern about, the collection of personal sociodemographic information by hospitals. We also examined public preference for practical approaches to the future collection of such information. RESULTS: In this sample of Canadian adults, nearly half did not believe it was important for hospitals to collect individual-level sociodemographic data. The majority had concerns that the collection of these data could negatively affect their or others' care; this was especially true among visible minorities and those who have experienced discrimination. There was substantial variation across participant subgroups in their comfort with the collection of various types of information, but greater discomfort in general for current household income, sexual orientation, and education background. There was consistent discomfort reported from older participants. Participants in general were most comfortable providing this type of information to their family physician. INTERPRETATION: The importance of collecting patient-level equity-relevant data is not widely appreciated in Canada, and our survey has shown that concern about how these data could be misused are high, especially among certain subgroups. Qualitative research to further explore and understand these concerns, patient education about data usage and privacy issues, and using the family doctor's office as a linked electronic data collection point, will likely be important as we move toward high-quality equity measurement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.027 | 0.062 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.007 |
| Science and technology studies | 0.007 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".