Taking Care of Business : The Clinical Research Data Edition
Bibliographic record
Abstract
Presentation for the Health Libraries Association of British Columbia. Take a moment to think about your researchers and the kinds of data they access and/or create. Where is this data stored and how is it organized? If they were asked to share data with another researcher would they be able to make sense of that work? If they needed to locate the data files from 5 years ago, how easy would it be to find and use datasets? What privacy and security controls have they implemented to ensure that they are appropriately safeguarding their data? Is their data governed by legislative requirements? If so, are they required to comply with additional policies and standards when using these datasets? If you are unsure about the answer to any of these questions you have come to the right place. This workshop will cover the basics of data management plans, metadata and data documentation, data privacy and security, and data sharing. It is designed to support researchers engaging in clinical research to start thinking about the steps they can take to better manage their research data. Kaitlyn Gutteridge is the Research Data Privacy and Security Officer for the ARC team. In addition, Kaitlyn serves as a member of the Compute Canada Security Council. She holds a Master of Science degree from the London School of Hygiene and Tropical Medicine where she focused her epidemiology training on multilevel modeling of chronic disease development. She previously held positions supervising the implementation of large-scale research initiatives at Simon Fraser University and the Centre for Hip Health and Mobility. Most recently, she served as the Privacy and Governance Lead at Population Data BC. In her position at Population Data BC, Kaitlyn served as the organization’s Privacy Officer and managed the negotiation, development, and execution of information sharing agreements and associated policies & procedures. Eugene Barsky is Research Data Librarian at the UBC Library. His recent peer-recognition included American Society for Engineering Education and Special Library Association awards. He published more than 20 peer-reviewed papers and presented at more than 40 conferences. Eugene is chairing the national Portage Data Discovery Expert Group, participates in building the Canadian Federated Research Data Repository (FRDR), and collaborates with Research Data Canada (RDC). Eugene is an adjunct faculty member at the iSchool at UBC, teaching courses in science librarianship and research data management, and is an active member of the Pacific Northwest data curators group.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.042 | 0.093 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.005 | 0.011 |
| Science and technology studies | 0.005 | 0.005 |
| Scholarly communication | 0.021 | 0.014 |
| Open science | 0.006 | 0.014 |
| Research integrity | 0.009 | 0.016 |
| Insufficient payload (model declined to judge) | 0.180 | 0.123 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".