MétaCan
Menu
Back to cohort
Record W2117288619 · doi:10.1136/bmj.332.7548.983

How should we rate research?

2006· editorial· en· W2117288619 on OpenAlexaboutno aff
Richard Hobbs, Paul M. Stewart

Bibliographic record

VenueBMJ · 2006
Typeeditorial
Languageen
FieldDecision Sciences
TopicEvaluation and Performance Assessment
Canadian institutionsnot available
Fundersnot available
KeywordsGovernment (linguistics)CornerstoneResearch Assessment ExerciseQuality (philosophy)InstitutionHigher educationPolitical sciencePublic relationsShadow (psychology)Medical educationBusinessPsychologyEconomicsMedicineEconomic growth

Abstract

fetched live from OpenAlex

Last month the UK chancellor, Gordon Brown, announced in his annual budget speech that the government intends to simplify the way it funds academic research. Pending a consultation exercise through 2006, the government wishes to replace the United Kingdom's unique research assessment exercise (RAE).1 The most radical proposal is to scrap peer assessment of the quality of each university's research, a cornerstone of the exercise, and to introduce assessment based mainly on metrics—effectively performance indicators. Possible metrics include research income, publications, citations, and numbers of research students, all of which correlated well with scores achieved in previous exercises.1 Preparations for RAE2008 are well underway and will proceed as planned, but the exercise will now incorporate a shadow system, using metrics alongside the panel based peer review system. It is now entirely possible that the allocation of research funding after 2008 will be based on metrics assessment rather than the historical peer review exercise.1 Coordinated by the UK's higher education funding councils, RAE2008 will be the sixth national peer evaluation of the quality of research conducted by higher education institutions.2 The exercise assesses the outputs of research by academic staff in higher education institutions. The rating for each institution is converted to a multiplication factor—which, alongside factors controlling the volume of research (such as staff numbers, numbers of research students, and peer reviewed grant expenditure), determines how much quality adjusted research funding the government gives to that institution. Along with money allocated for teaching, this funding represents the education councils' block grant to UK higher education institutions; for 2007-8 the higher education funding council will allocate £1.45bn (€2.1bn; $2.6bn) to institutions as quality adjusted research funding. This research quality assessment exercise is the largest attempted anywhere in the world. Most countries continue to base their funding of public research on negotiation with academic institutions (Austria, France) or on simple formulaic models factoring in numbers of staff and students (United States, Canada, most of Europe). Only Hong Kong replicates the intensity of the research assessment exercise, though some countries such as New Zealand intend to adopt similar evaluations. In Australia, quality evaluation based on performance indicators is probably the closest system to the new proposed metrics analyses.3 Evidence suggests that the research assessment exercise is having a beneficial impact on the quality and competitiveness of UK research. It has generated a cycle of quality rating: institutions rated highly receive more money and are able to do more research of high quality. This has concentrated investment in research in the UK; the proportion of total public funding to the top tenth of researchers in the UK increased from 47% in 1980-1 to 57% in 1997-8, compared with a decline in the US from 47% to 43%.4 The exercise has increased the quality of research perhaps through focusing the research strategies of academic institutions and providing incentives.4 In RAE1996, 32% of staff worked in units rated as excellent, and this rose to 55% in RAE2001. Overall, UK research tops the world league for papers per dollar expended and citations per dollar expended, and it is fourth ranked for papers per researcher.5 Maintaining research investment will remain essential if the UK is to retain its international profile. Delivering 9% of the world's research effort for 4.5% of the world's research expenditure in a high cost economy seems precarious, but the research assessment exercise has provided an evidence base to justify such an investment. The results of the exercise are also used for many secondary purposes. For example, each institution's rating may influence the chance of receiving other research funding or of fellowships or studentships being awarded. Most research sponsors now request applicants to state their institution's score, and eligibility for some initiatives is limited to researchers from institutions with the highest ratings from the most recent assessment (in 2001). RAE2008 represents the next cycle in the process but with several notable differences from earlier exercises. Staff engaged in research who are in post in the institution on the 31 October 2007 census date are invited to submit up to four research outputs published between the dates of 1 January 2001 and 31 December 2007; most are expected to be publications in peer reviewed journals, including systematic and Cochrane reviews.2 Metric data is certainly required—each institution must provide, for the period from 1 January 2001 to 31 July 2007, data on research studentships and associates and expenditure through research grants. In contrast to earlier exercises, however, and somewhat at odds with the “metric” lobby, a far greater emphasis is placed on peer review of research outputs in generating an overall score; 75% of the total score in any given assessment will be derived from publication outputs,2 with the scoring system placing more emphasis on research of international quality. Instead of the RAE2001 rating of 0 to 5*, RAE2008 will use an overall quality profile assessment of “unclassified” (below national standard), “one star” (nationally recognised research), “two star” (internationally recognised research), “three star” (internationally excellent research), and “four star” (world-leading research). In the absence of any published criteria, peer reviewers might find it difficult, however, to rate international research using the subjective concepts “recognised,” “excellent,” and “world-leading.” Additional assessment panels have been created to provide more emphasis on important research themes such as cardiovascular disease, cancer, and neurosciences while also recognising the methodological diversity of the more applied research areas of clinical epidemiology, health services research, and primary care. These latter areas were rated below laboratory based and hospital based research in previous research assessment exercises but are particularly important to the NHS, for instance in guiding policy, identifying cost effective ways of configuring care, and providing high quality syntheses of research evidence. In an era of perceived crisis in training academic clinicians,6,7 RAE2008 will be both novel and innovative in encouraging institutions to submit younger researchers (clinical lecturers, for example) and research collaborators funded in adjacent NHS trusts; in both cases only two research outputs will be required. Accompanying text will allow institutions to outline their research strategy and vibrancy of research training programmes.2 The parallel, metric based analysis will fail to take into account many of these important additions to RAE2008. The cost and the workload of the exercise will undoubtedly be further drivers down a metrics route. For RAE2001 the direct costs were over £5m, an increase of 68% above costs in RAE1996.8 The amount of academic time spent was illustrated by 2598 submissions from 173 institutions, most listing four papers from each of 48 022 researchers.5 However, despite this direct expenditure, with estimates for institutional and opportunity costs inflating the 2001 figure to over £37m, the overall cost of the research assessment exercise represented just 0.8% of the total funding allocated on the basis of the exercise.9 Money well spent, or an unnecessary burden when faced with a metrics alternative? The government has announced its intent to instigate change down this route; already, related metric based approaches are being used to guide allocation of major NHS research and development funding, for example.10 Whatever the pitfalls, our entire academic base is based on the philosophy underpinning peer review and it will be of paramount importance in RAE2008 to compare the peer review and metric based systems before recommending major changes. Equally, RAE2008 should inform the international community on the optimal approach to assessing research quality.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.039
metaresearch head score (Gemma)0.021
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Scholarly communication, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Insufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Editorial · Consensus signal: Editorial
Teacher disagreement score0.262
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0390.021
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0020.000
Open science0.0020.000
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.550
GPT teacher head0.625
Teacher spread0.075 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEditorial

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations18
Published2006
Admission routes1
Has abstractyes

Explore more

Same venueBMJSame topicEvaluation and Performance AssessmentFrench-language works237,207