Have research assessment exercises improved the quality of nursing research?
Notice bibliographique
Résumé
Do government research quality assessment exercises help or hinder the growth of nursing and its research? Many academics in Europe, Asia and Australasia will be familiar with these performance-based exercises that, for the last 20 years, have been used by governments to assess research quality and determine between 2–25% of institutional and departmental funding (Martin 2011, Hicks 2012). The assessment of research quality dates back to the 1970s when it was first used to explore the value of science to society (Martin 2011). During the mid-1980s, the assessment of scientific impact became more formalized and was used then in the UK to make funding decisions (Martin 2011). More recently, there has been a concerted shift towards assessing the broader societal impact and reach of research (Bornmann 2013), including commercialization (Martin 2011). The value and validity of metrics in assessing such aspects – including research quality and influence – is coming under increasingly intense scrutiny (Wilsdon et al. 2015). Currently, over 14 such assessment systems now exist internationally (Hicks 2012) which share a common focus on: concentrating resources, fostering international publication (in English) and incentivizing excellence (Hicks 2012). Across higher education institutions, both old and new, there is now widespread recognition from staff that universities are businesses that must seek and maintain income from teaching and research (Kok et al. 2010). Maximizing performance in relation to the exercises is then widely seen as important. However, the research exercises can make or break funding and the reputations of individual researchers, their departments and leaders. While nursing mostly has its own dedicated stream of assessment as a relatively young academic discipline and presence in universities, have these exercises influenced the growth of the quality and visibility of nursing research? What are the harms and costs of such exercises to the state of today's nursing research? Examining the likely effects of the exercise at the disciplinary, department and individual level provides limited assurance or positives for nursing. In terms of the discipline, there is little evidence that the mere presence of research assessments has improved quality because the relationships between assessment and quality are notoriously problematic. While few people argue that the impact of research is unimportant, many argue about whether and how to do this (Molzahn & Clark 2015). Impact can be both positive and negative (Martin 2011), usually requiring a wide range of measures (Eisen et al. 2013), and, even then, may have an indirect and complex relationship with research quality (Martin 2011). It remains doubtful that merit can be reduced to a single measurable quality of a paper because such a wide variety of factors influence citation rates and these practices vary across disciplines (Eisen et al. 2013). The relatively low enduring impact of most nursing journals, while troublesome, has mostly contributed to problems with scientific reputation than scientific quality. This follows because reviewers of academic papers in other disciplines tend to assume wrongly that lower impact and/or lower cited papers equate to lower quality (Eyre-Walker & Stoletzki 2013). Although journal impact scores in nursing remain low and fairly static, the effects of this and the mere presence of assessment initiatives on the quality of nursing research has likely been modest. At the departmental level, while assessment exercises are intended to assess quality, as with most such performance assessment systems, research activities and strategies evolve to work towards the assessments. Seen variously as ‘game-playing’ (Martin 2011) or thoughtful opportunism, this further questions whether the assessments primarily assess and drive quality. For individual nurse researchers, assessment exercises in some countries act to foster a greater free market and higher competition for recruitment. While this may have led to short-term hiring and increased salaries for some individuals, it leaves long term development for individuals and departments less to the fore. Individuals can be hired solely based on their ability to boost short-term performance rather than for longer term fit of both personality and work. Collectively, these concerns have led commentators to ‘raise fundamental questions about how we currently evaluate science and how we should do so in future…current research assessment practice is neither consistent nor reliable’ (Eisen et al. 2013 p.1). It is easy to hate the exercises – what they involve and represent – and many do. Reducing the influence of one's research to numbers remains a common resentment for researchers across disciplines (Wilsdon et al. 2015). By necessity, this filters and reduces nuance. Recoding, collation and submission takes extensive amounts of staff time and energy to determine strategy, collate submissions and implement. Stakes and stress for those involved are high: jobs are often literally ‘on the line’. The demands of submission mostly act to reduce the amounts of time and energy actually devoted to the research the exercises seek to evaluate. The exercises represent a form of external and even governmental influence over academic work and its impact and publication. While many peer academics are involved in the actual reviews, governments still make the rules around what counts and how. This external control can be seen less as a factor to be triangulated with academic freedom than a direct threat to academic freedom. In submitting departments, the submissions reduce years of complex research from many people to a single summative measure of quality. They pit departments, often with contrasting research profiles, against each other and decide the fate of individuals in the same department as being ‘in or out’ of submissions. Suddenly, an individual's own CV becomes a platform for considering competitiveness not just in disciplinary terms but in relation to new and varied indicators, such as journal impact factors, categorizations and number of completed PhD students. Yet, we also believe that it is important to see potential good arising from the assessments – not only to accept the reality of their existence but also to use it to promote positive developments for the discipline of nursing, for our departments and for us as individual nurse academics. For the discipline, recent research assessment initiatives often require and incentivize academics to work in non-traditional ways – to publish more ambitiously and consider the broader impact of their work on society beyond publication. For example, this past iteration of the UK Research Excellence Framework required researchers to describe the impact of their work less in terms of academic impact. The measurement of impact has moved from purely measuring scientific impact to a broader sense of societal impact (Bornmann 2013). While nursing has little track record of contributing to some of the more recent measures of research impact – like patents and spin off companies – nurse researchers frequently claim to be doing work that matters to patients, their families and communities and seeks to more directly affect the world. Researchers tend to cite this patient-good frequently as a motivation for their research and often do seek for research to have more direct clinical benefits – a possibility due to the nature of nursing – that is extremely realistic. Nevertheless, whether this occurs frequently or as frequently as these claims would indicate is far more debatable. The exercises compel departments and researchers to consider such patient and societal impacts and nursing is very well placed to do work that makes a real and lasting impact in these wider ways. At the department level, nursing has had the benefit of frequently having its own stream of assessment in these exercises. For example, in both recent national research assessments in Australia and the UK, nursing had its own separate review panels, assessments and results. As such, academic nursing departments competed directly against each other for proportions of funding rather than in more ‘generic’ pools of various health sciences disciplines. This separation offers arguably fairer comparisons given the relatively lower impact of nursing journals and recognizes that nursing is less established in higher education settings than psychology or public health. However, it implies that nursing should be treated differently than other health disciplines – a recognition out of step with many countries where nursing has to compete with those from other disciplines for research funding and publication in journals. This ‘special status’ is also at odds with the predominant trends towards interdisciplinary research teams which encourage collaboration and publication across traditional disciplinary boundaries. It also continues to risk re-emphasizing that as an academic discipline, nursing somehow is different than other applied evidence-informed health disciplines and should be held to different, often weaker, standards. While some may be sceptical that nursing still asks for and holds ‘special status’, the continuingly disappointing performance of the nursing professoriate is unlikely to be improved by continued isolation (Thompson & Darbyshire 2013). Without the presence of such exercises (as is the case in Canada), departments and their staff have extensive autonomy but limited financial or ‘real’ incentives to encourage academics to publish with aspiration, recruit graduate students, or bring research funding in. Rather than leading to a utopian freedom and success through academic liberation, this mostly leads a lack of scholarly vision, strategy and urgency in individuals and departments. These are catastrophic to the success of nursing as a fledging discipline. Too often individual academics in such countries publish for quantity over quality and visibility and departments fail to recruit strategically and end up with no discernable focused areas of research to concentrate expertise or differentiate themselves from their peers. In the absence of any comparable data, individuals and departments frequently end up with a falsely superior sense of their own success compared with peers. Departments become collections of individuals pulling in every direction rather than units that recognize their individual and collective performance are intertwined. Facing the future, research quality assessment exercises are here to stay. As we have shown, they offer challenges, opportunities, and threats for nursing. They expose vulnerabilities in our discipline's research progress to date, test our abilities to raise our individual and collective research game, and give nursing high profile chances to demonstrate the differences nurse research claims to make. With increasingly sophisticated metrics and broader impact assessment, it is increasingly untenable to state that nursing research benefits patients and practice while also undermining data-driven attempts to evaluate such impact. Nursing research's aims and the assessment games can indeed be good for each other.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,102 | 0,048 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,001 | 0,002 |
| Études des sciences et des technologies | 0,003 | 0,002 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,002 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,015 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».