MétaCan
Menu
Back to cohort
Record W3209033102 · doi:10.5281/zenodo.3543505

Investigating the Link Between Research Data and Impact

2019· article· en· W3209033102 on OpenAlexaff
Eric Jensen, Mark Reed

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2019
Typearticle
Languageen
FieldComputer Science
TopicResearch Data Management Practices
Canadian institutionsImpact
Fundersnot available
KeywordsLink (geometry)Computer scienceComputer network

Abstract

fetched live from OpenAlex

The Institute for Methods Innovation – a research charity registered in the United States and United Kingdom – was commissioned by the Australian Research Data Commons (ARDC) to investigate how research data contributes to non-academic impacts, drawing on existing impact case studies from the UK Research Excellence Framework. Project overview The research involved analysing impact cases from the UK’s Research Excellence Framework (REF). These cases were sifted to only review high scoring cases with a strong emphasis on ‘data’. Relevant text to this research was extracted from the larger impact narratives. A content analysis was conducted to identify patterns, linking research data and impact in the narratives. This analysis achieved a high level of reliability, based on established methodological standards. What type of impact was developed from research data? The most prevalent type of research data-driven impact related to Practice (45%). This category of impact includes changing the ways that professionals operate, changing organizational culture and improving workplace productivity or outcomes. It also includes improving the quality of products or services through better methods, technology, understanding of the problems, etc. Government impacts were the next most prevalent category identified in this research (21%). This category includes reducing the cost to deliver government services, enhancing the effectiveness or efficiency of government services and operations and providing input into government planning, decision-making and policymaking. Other relatively common types of research data-driven impacts were Economic impact (13%) and General Public Awareness impacts (10%). How was impact developed from research data? Impact from research data was developed most frequently through Improved Institutional Processes / Methods (40%). This relates to making an institution’s way of operating better, more efficient or effective at delivering outcomes. The second most common way of developing impact was via a report (32%) of some kind, that is, pre-analysed or curated information. Analytic Software or Methods (26%) comprised the third most frequently used way of developing impact. Here, research data are used to generate or refine software or research and analytic methods. Who benefited from the research data-linked impact? Professionals (50%), Government, Policy, or Policymakers (42%) and Industry / Business (38%) were the most common types of beneficiaries from the research data-linked impact. This finding is partly explained by a two-step flow of research data-linked impact that ultimately reaches publics or wider non-academic stakeholders. While intermediaries such as professionals, policymakers and industry are primary beneficiaries or users of the research data-based impact, they in turn use what they have gained to develop insights, services, products and policies that deliver broader public impacts. Looking at patterns in this analysis, the following correlations were identified: Searchable databases tended to be used with the general public (r = .22), while ‘enhancing institutional processes / methods’ is not (r = -.26). Analytic software (r = .23) and ‘improved institutional processes / methods’ (r = .32) were used more to develop impact with industry / business. Sharing of raw data was more often an impact development pathway with environmental impacts (r = .2) than other types. Conclusions The analysis found that research data on their own rarely develop impact, but instead they require analysis, curation, product development or other strong interventions to leverage broader non-academic value from the research data. These interventions help to bridge the gap between research data- which might otherwise go unused for the purpose of developing impact- and the diverse range of potential primary and secondary beneficiaries. In the same sense, the impact of research data can be engineered, through closer links between government, industry and researchers, capacity building for researchers to effectively use research data to develop impact and capacity building for potential beneficiaries to establish links with researchers and to access and make sense of useful sources of research data that can be adapted to serve new purposes. Moreover, the way that research data is made available, and the nature of the support available, can affect how feasible it is to use that research data to develop new and creative pathways to impact. Finally, there were surprisingly high ‘uniqueness’ scores for the impacts linked to research data (97%), suggesting that most of the research-data linked REF-reported impacts may have only been possible to develop through research data. However, limitations inherent in REF impact case studies have to be taken into account before drawing firm conclusions on this point. The Dataset 2_ARDC - Analysis Data.csv : Core dataset 3_ARDC - list of cases : List of all REF cases used in the analysis 4_ARDC - list of variables : Breakdown of all coding variables and values. Refer to the Coding Guide for a detailed description of each code. 5_ARDC - ICR Data : Inter-coder reliability dataset Other Resources For text mining UK REF Impact Case Studies a collection of R scripts is available on GitHub

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.497
metaresearch head score (Gemma)0.804
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.503
Threshold uncertainty score0.620

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4970.804
Meta-epidemiology (narrow)0.0010.003
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0250.034
Science and technology studies0.0070.021
Scholarly communication0.0260.035
Open science0.0040.034
Research integrity0.0040.008
Insufficient payload (model declined to judge)0.0090.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.223
GPT teacher head0.377
Teacher spread0.154 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designObservational
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)Same topicResearch Data Management PracticesFrench-language works237,207