BC Data ScoutTM: A New Tool To Investigate Datasets For Health Research
Bibliographic record
Abstract
IntroductionBC’s Ministry of Health (MOH) maintains many administrative databases with rich information and analytical potential. Researchers are keen to use these data for both discovery and applied research. Historically, limited views of data availability and populations therein have supported study feasibility. Therefore, we developed BC Data ScoutTM, a cohort browser. Objectives and ApproachWe developed a cohort browser service to provide information to researchers planning a study using MOH data. The objective was to create a tool that is simple to use, provides quick results and is free to users to encourage its use. A better understanding of the data available can improve study quality and expand the user-base by giving researchers access to information not previously available during the planning stages. The tool will be evaluated by examining the number of requests received and a user satisfaction survey. Plans are in place to expand into additional data sources and extend query sophistication. ResultsThe BC Data ScoutTM online tool provides cohort information in the form of highly aggregated, approximate results to researchers planning a study. It was developed by the MOH, the BC SUPPORT Unit and Population Data BC (PopData) and was launched in February 2018. The service is delivered by PopData. BC Data ScoutTM offers province-wide information for query, is accessible to a wide group of eligible researchers, and has data availability from the year 2000 onwards. Four types of MOH data are available for query: hospital data; physician data; pharmaceutical data; and demographics. In addition to determining study feasibility, the aggregate reports also help to further refine a full data access request and provide enough information to complete and strengthen a funding application. Conclusion/ImplicationsBC Data ScoutTM will be beneficial for researchers planning to request data. This preliminary information may increase the chances of meaningful research studies to obtain funding, and the production of relevant, high-quality research results. BC will be among the first jurisdictions across Canada to offer this type of feasibility service.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.051 | 0.164 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.013 | 0.015 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.008 | 0.011 |
| Open science | 0.004 | 0.013 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.102 | 0.025 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".