Empowering Immigrant Community through Ownership, Control, Possession and Utilization of Data: Community Based Health Data Cooperative
Bibliographic record
Abstract
Background: Canadian immigrant populations come from diverse ethno-geographical backgrounds and exhibit differences in their culture and understanding of health and wellness. It is imperative to collaborate and empower these diverse communities. This can be achieved through establishing a health data cooperative (HDC) that allows the availability of the valuable data for societal purposes. HDC is a health data bank where cooperative members collect, store, use and share healthrelated data (e.g., health condition, lab results, social determinants of health data, etc.). The aim of this review is to analyse the feasibility of immigrant based HDC model through conducting stakeholder and customer discovery interviews to gain insights and conduct early-stage assessments determining sustainability and scalability of the HDC. Methods: We propose to undertake a comprehensive environmental scan including stakeholder analysis to conduct key-informant interviews. Through the interviews, we will gather feedback and research existing models of cooperatives to develop the HDC framework. Expected results: Through this comprehensive environmental scan, we are hoping to engage stakeholders and explore key components such as ethical & legal frameworks, organization management, data security, privacy, computing science, knowledge mobilization, and community development. We believe by enabling the immigrant communities has the potential to promote health equity through empowerment (enhance individual competence & self-esteem, increase community action & participatory learning exercises). Conclusions: The results of this project are an informative first step to launch a pilot HDC model in immigrant communities. Moreover, further research on scalability and performance would be required as the HDC model becomes operational.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.016 | 0.007 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.002 | 0.013 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".