Research Data Management – Approaches to Capacity Building by Acting Locally While Thinking Nationally
Bibliographic record
Abstract
The role of university libraries in research data stewardship has been in rapid growth and evolution.Key principles for good research data management standards are emerging and stabilizing internationally providing an opportunity for institutions to encourage and facilitate sound research data management practices among its students and researchers.Libraries must consider how we strengthen our collective ability to anticipate and respond to these needs.Using recent theoretical models of research data management services, this paper looks at an approach to build capacity around research data management services at a local level (the University of Ottawa Library) against the backdrop of a maturing national initiative (the Canadian Association of Research Libraries' (CARL) Portage Network).Using a theoretical framework consisting of OCLC's Tour of the Research Data Management (RDM) Service Space [Bryant, Lavoie, Malpas & OCLC, 2017], COAR's Librarians' Competencies Profile for Research Data Management [Schmidt & Shearer, 2016] and Cox, Kennen, Lyon & Pinfield's (2016) model of research data service maturity, this article will examine how a single institution, the University of Ottawa, is progressing in developing its own RDM service offering against the backdrop of a maturing national initiative, Portage, led by the Canadian Association of Research Libraries (CARL). Frameworks for RDM services and competenciesResearch data management, defined at its simplest, can be described as "the effective handling of information that is created in the course of research" [JISC, 2016].Recognizing RDM as a significant development requiring high levels of engagement by academic libraries, research library associations, such as the Association of Research Libraries (ARL), OCLC and the Canadian Association of Research Libraries (CARL) have been actively engaged in these areas, in response to interest from their member institutions.The benefits of good data management, from the perspective of researchers, administrative and funding agencies, and journal publishers, are well documented in the literature.Important issues such as researcher efficiency, re-use of publicly funded research data to create new investigational possibilities, and the support of published articles by sharing data which underpin conclusions, are driving the development of research data strategies on campus.These benefits, along with other factors, contribute to the rising interest and need for robust discussions around each institution's unique profile of priorities, research activities, funding, policies, capabilities and aspirations.More recent models of RDM services place a greater emphasis on consultation services and access to infrastructure, illustrating how our shared understanding of RDM services is evolving and, in the case of some institutions, has begun to shift from being a strategic issue towards more of a procedural issue [Pinfield, Cox & Smith, 2014].In March 2017, OCLC published the first of a four report series looking at the research data management service space [Bryant et al., 2017] .Reviewing services developed by over a dozen research libraries across three continents, the authors propose a model representing three categories of RDM services commonly found throughout academic libraries.These categories include (i) education service, (ii) expertise service, and (iii) curation service.In 2016, COAR's Joint Task Force on Librarians' Competencies in Support of EResearch and Scholarly Communication published a set of competencies [Schmidt & Shearer, 2016] describing the skills and knowledge needed by academic librarians to support researchers in managing their research data.Schmidt and Shearer define three categories of RDM support that may be offered by libraries, these include (i) providing access to data, (ii) awareness and support for managing data, and (iii) managing a data collection.Both the OCLC and COAR reports describe RDM services being offered at research libraries around the world.There is considerable overlap in the types of services and activities described in both reports, with instruction along with services to build awareness for managing data, most frequently mentioned.According to both studies, many of the roles that the library may increasingly play can be categorized under OCLC's Expertise Services, indicating that there is a distinct need for customized assistance in support of individual research projects.Schmidt and Shearer touch lightly on the role of the library in technical infrastructure support and development, activities largely found under the Curation Services category of the OCLC report.Appendix A provides examples of services organized by each OCLC service category.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.111 | 0.101 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.010 |
| Science and technology studies | 0.019 | 0.062 |
| Scholarly communication | 0.043 | 0.053 |
| Open science | 0.010 | 0.038 |
| Research integrity | 0.008 | 0.010 |
| Insufficient payload (model declined to judge) | 0.014 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".