Mapping Canadian Data Assets to Generate Real-World Evidence: Lessons Learned from Canadian Real-World Evidence for Value of Cancer Drugs (CanREValue) Collaboration’s RWE Data Working Group
Bibliographic record
Abstract
Canadian provinces routinely collect patient-level data for administrative purposes. These real-world data (RWD) can be used to generate real-world evidence (RWE) to inform clinical care and healthcare policy. The CanREValue Collaboration is developing a framework for the use of RWE in cancer drug funding decisions. A Data Working Group (WG) was established to identify data assets across Canada for generating RWE of oncology drugs. The mapping exercise was conducted using an iterative scan with informant surveys and teleconference. Data experts from ten provinces convened for a total of three teleconferences and two in-person meetings from March 2018 to September 2019. Following each meeting, surveys were developed and shared with the data experts which focused on identifying databases and data elements, as well as a feasibility assessment of conducting RWE studies using existing data elements and resources. Survey responses were compiled into an interim data report, which was used for public stakeholder consultation. The feedback from the public consultation was used to update the interim data report. We found that databases required to conduct real-world studies are often held by multiple different data custodians. Ninety-seven databases were identified across Canada. Provinces held on average 9 distinct databases (range: 8-11). An Essential RWD Table was compiled that contains data elements that are necessary, at a minimal, to conduct an RWE study. An Expanded RWD Table that contains a more comprehensive list of potentially relevant data elements was also compiled and the availabilities of these data elements were mapped. While most provinces have data on patient demographics (e.g., age, sex) and cancer-related variables (e.g., morphology, topography), the availability and linkability of data on cancer treatment, clinical characteristics (e.g., morphology and topography), and drug costs vary among provinces. Based on current resources, data availability, and access processes, data experts in most provinces noted that more than 12 months would be required to complete an RWE study. The CanREValue Collaboration's Data WG identified key data holdings, access considerations, as well as gaps in oncology treatment-specific data. This data catalogue can be used to facilitate future oncology-specific RWE analyses across Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.262 | 0.475 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.025 | 0.039 |
| Science and technology studies | 0.011 | 0.007 |
| Scholarly communication | 0.022 | 0.011 |
| Open science | 0.011 | 0.015 |
| Research integrity | 0.004 | 0.009 |
| Insufficient payload (model declined to judge) | 0.009 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".