Bibliographic record
Abstract
Data and its computation have become daily news topics. The Globe and Mail February 3 article on Canada’s inadequate big data computing capacity, and a recent announcement by the International Committee of Medical Journal Editors (ICMJE) proposing that all authors share the de-identified individual-patient data underlying the results in their articles, are only 2 of many examples. The editors of the New England Journal of Medicine responded to ICMJE with an editorial on "research parasites," stating that the culture of data sharing is far from being universally embraced. Underlying this news is the importance of research data management and the importance of Research Libraries in implementing Research Data Management strategies. Chuck will speak about developments in research data management services and infrastructure across Canada and about the support Portage will provide Canadian higher education institutions. He will discuss collaborative initiatives between the library community and other research stakeholders, including the January 27th announced memorandum of agreement with Compute Canada and emerging data policies from funding agencies. Time will be included for a good discussion. Portage aims to coordinate and expand existing library-based expertise, services and infrastructure so that Canadian researchers will have access to the support they need for research data management (RDM). Portage will have two major components: a network of expertise to provide access in both English and French to a comprehensive set of resources, tools and experts; and a preservation and discovery system to connect the various infrastructure and service components needed for national preservation and discovery of data. Chuck Humphrey has supported data services at the University of Alberta since 1992, and has worked on numerous regional, national and international initiatives to increase access to data for teaching and research purposes. He was involved in OECD Global Science Forums on Data and Research Infrastructure for the Social Sciences in 2010-2011 and on Ethics and Big Data in 2014-2015. Chuck was the lead investigator on a University of Alberta Libraries’ successful application for a data centre in the Canadian International Polar Year (IPY) Data Assembly Centres Network, which has now become the Canadian Polar Data Network. He currently serves on the Steering Committee of Research Data Canada; is a Board member of the Consortia Advancing Standards in Research Administrative Information; and has been a key participant in CARL’s Project ARC working group, which developed the vision and framework for Portage.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.336 | 0.554 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.006 | 0.004 |
| Bibliometrics | 0.020 | 0.019 |
| Science and technology studies | 0.010 | 0.026 |
| Scholarly communication | 0.062 | 0.045 |
| Open science | 0.014 | 0.071 |
| Research integrity | 0.012 | 0.026 |
| Insufficient payload (model declined to judge) | 0.019 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".