MétaCan
Menu
Back to cohort

The British Association of Dermatologists therapeutic guidelines: can we AGREE?

2003· editorial· en· W2166018121 on OpenAlexaboutno aff
N.H. Cox, Hywel C Williams

Bibliographic record

VenueBritish Journal of Dermatology · 2003
Typeeditorial
Languageen
FieldMedicine
TopicClinical practice guidelines implementation
Canadian institutionsnot available
Fundersnot available
KeywordsMedicineQuality (philosophy)Family medicineAlternative medicineRandomized controlled trialMEDLINEMedical educationPolitical scienceLawPathology

Abstract

fetched live from OpenAlex

The British Association of Dermatologists (BAD) started a process of producing guidelines for dermatologists 5 years ago. Since then, 10 guidelines have been published or are in press, and a further four have completed a consultation period (Table 1). Several additional topics are in progress, also listed in Table 1. Strengths of the BAD guidelines include the fact that the development process is overseen by a committee, most of whom are not the authors, and also that the entire membership of the BAD has an opportunity to comment on a draft version (or even to apply the draft recommendations) for a 3-month consultation period. Most of the guidelines include practising district general hospital representation in the authorship, are written with practical application in mind, and are specifically targeted at dermatologists in secondary care. Writing the guidelines has required a pragmatic approach such that the recommendations take account of the availability of treatments within the U.K. National Health Service, and there has also been a commitment to make recommendations even in situations where high-quality evidence is not available (a problem that applies to many older ‘established’ therapeutic approaches). Guidelines are different from systematic reviews of evidence of treatment efficacy based solely on randomized controlled trials: although guidelines should be based on high-quality evidence if it is available, they can still cover areas where evidence is scanty or of poor quality if indeed the real world requires such guidelines. However, over the last 10 years, and certainly in the last 5 years since our process started, the whole topic of clinical guidelines has turned into a new industry. In particular, consideration now has to be given to the many ‘guidelines for guidelines’ that abound. The purpose of this Editorial is to comment on two major issues that will affect the BAD guidelines process, and to update readers on how the Therapy Guidelines and Audit Subcommittee of the BAD will address these. The two specific issues that will be addressed are the use of tools for validation of guidelines (also referred to as guidelines evaluation or appraisal), and developments in grading of evidence and linking with management recommendations. Numerous tools have been used to validate guidelines. There are various items that should ideally be incorporated into any guidelines (for example, see Table 2). Assessment of whether these items have been included can be scored in various ways with validation tools. However, even the validation tools have potential areas for improvement. One study examined 15 of these, containing 8 to 142 questions about a series of attributes; none of the validation tools covered all 44 topics considered by the authors to be important for guidelines content analysis (the range was from 6 to 34 items covered), and the study therefore concluded that all of the validation tools examined had some failings.1 Both guidelines and validation tools are probably still on a learning curve. We will concentrate on the Appraisal of Guidelines for Research and Evaluation (AGREE) instrument,2 as this guidelines validation tool scored well in the study discussed (28 attributes covered), and has been adopted by the Clinical Effectiveness and Evaluation Unit of the Royal College of Physicians for ‘scoring’ guidelines that are listed on its database.3 This validation tool includes 23 questions divided among six domains (briefly listed in Table 2); each is scored on a four-point scale by a series of appraisers in order to reach a validation score. However, there are specific problems with some of these; for example, a conflict of interest statement is either included or it is not, so the intermediate scores are redundant, and the scoring process uses an averaging process in which outlying scores may distort the assessment. More importantly, guidelines users will generally be most interested in the overall content, usefulness and applicability of the guidelines, which are only a minor part of the final score. There are various ways in which guidelines may perform suboptimally against these criteria, some of which are more to do with a lack of explicit statements rather than poor retrieval of the evidence or incorrect inferences. In general, we would expect the ‘Scope and purpose’ domain of the BAD disease management guidelines to be fairly strong, as outlined in Table 2, and there are a number of areas where we should consistently fulfil objectives listed in the AGREE document, for example: (i) all of our guidelines are identified as being ‘for dermatologists’ so there is no doubt about the target users; (ii) we have increasingly used lists and tables for the main recommendations, have provided summaries and lists of possible audit points, and have supplied appendices listing the grading of evidence and strength of recommendation criteria, all of which are suggested attributes; and (iii) many of the guidelines do attempt to balance risks and benefits, and to take into account different situations; for example, the size and site of lesions in patients with Bowen's disease,4 the thickness of melanoma5 or the histological type of basal cell carcinoma.6 There are, however, some areas where the BAD guidelines are likely to score low against the AGREE instrument. For example, we have no formal piloting process, and we have not included patients in the development process (although we do have nondermatological representation on the overseeing committee). One can argue that the patient perspective may be more important for some guidelines than others; for example, the relevance might be low for a one-off severe disorder such as toxic epidermal necrolysis but might be much greater for a chronic skin condition such as atopic dermatitis. However, most of our potential areas of low scoring can easily be resolved through a more explicit process. The issues that need to be addressed fall into three main categories: There are some important pieces of information that have been omitted from the guidelines to date. For example, a conflict of interest and funding statement has not routinely been included. This may sound trivial, but a recent study evaluating the relationship between authors of clinical practice guidelines and the pharmaceutical industry found that 57% of guidelines authors had some form of interaction with the drug industry and that most of the guidelines they were involved with did not require a formal process for declaring these relationships.7 Only 7% of these guidelines authors felt that their own relationship with pharmaceutical companies influenced their recommendations, although 19% felt that their coauthors' recommendations were influenced by their relationships. As this Editorial will become a standard part of our background documentation for the guidelines process, let us at this point make it clear that no funding is specifically linked to any BAD guidelines, that authors and the committee are not paid for writing and commenting on the guidelines, and that the only external funding is a donation from the British National Formulary for dissemination of the guidelines summaries that are discussed below. The BAD already has a conflict of interest form for committee members but the Therapy Guidelines and Audit Subcommittee of the BAD will also need to ensure that all authors acknowledge potential conflicts of interest and that these will be published openly in the body of the guidelines. Data relevant to the guidelines development process have been published elsewhere but not in the text of the guidelines. The most important example of this is the summary of the BAD guidelines process published with the first of our guidelines in 1999.8 For example, this article described the 3-month consultation period with the entire membership of the BAD (‘External review’ question in the AGREE ‘Rigour of development’ domain). This has been included in the references list of the more recent guidelines so it is available to any guidelines appraiser (provided they read the guidelines carefully enough!) and it is documented in the appendix to each guideline as background data. Repeating it in detail in each guideline would increase the length of the article unacceptably. Even since the original summary document, the process has altered; in particular, there has often been an expansion of the original proposed three-author process and there has been collaboration with other specialist societies (for example, on melanoma,5 squamous cell carcinoma9 and lichen sclerosus10). Planned changes are already scheduled but are not yet available. Examples of these are possible changes to the links between level of evidence and grading of recommendation, discussed below, and also the recommendation for the ‘provision of additional materials’ that we have in hand: members of the BAD will shortly receive laminated single-page summary sheets of guidelines for easy access to the main guideline recommendations. Validation tools suggest that guidelines should have an expiry date, but it is already apparent that some aspects alter little while others are subject to rapidly changing literature and require early review. We will therefore aim to publicize any updates on the BAD website at clinically relevant points rather than at predetermined times from publication. The single-sheet summaries will carry some updates as well. Details of literature search criteria, such as the databases examined or use of existing Cochrane reviews, can readily be added to the text of forthcoming guidelines. Overall, within certain constraints (such as the fact that our BAD guidelines are written for a specific narrow target audience), we believe that we can achieve most of the issues listed in the AGREE guidelines assessment tool. Factors which will be common to all our guidelines are listed in Table 2, and this article is intended as a specific background document that accompanies our guidelines but will not be published within the text of each. However, in using the AGREE scores it should be borne in mind that these cover a whole domain (Table 2) and are related to a process: they do not necessarily reflect the strength of clinically relevant content or the utility of the guidelines in a practical situation. Turning therefore to the grading or level of evidence and the extrapolation to providing a ‘strength of recommendation’ conclusion, we are again faced with a plethora of newly developed options. Before discussing these, it is worth pointing out that grading the level of evidence and making recommendations based on this are two slightly different but related issues. Grading the quality of existing evidence, as its name implies, involves making a decision on the quality of evidence related to its tendency to minimize bias. Thus randomized controlled trials are generally considered to be a more robust form of evidence than are case series. Deciding the strength of a recommendation is a little more tricky as this depends so much on the clinical context of the patient to which it are being applied. As a result, some working in the guidelines field have dropped making recommendations for practice situations altogether and instead just grade the external evidence. This is the position adopted by one of the longest standing evidence-based medicine guidelines groups,11 which has produced over 400 guidelines for primary care practitioners in Finland based largely on Cochrane reviews. One of the problems is that complex human decisions on individual patients cannot be captured using simple categories, although some recent approaches using fuzzy set theory may turn out to be more appropriate.12 The current BAD evidence grading system is based on the 1979 Canadian Task Force recommendations.13 Based on feedback from our guidelines users, these Task Force recommendations are now viewed as being a bit vague especially in the middle grades of evidence, and at times not providing a robust link between evidence and recommendations. The problem is that, at the other end of the scale, grading systems may be very complex. For example, the Oxford Centre for Evidence-based Medicine Levels of Evidence manual uses 10 levels of evidence, which require understanding of homogeneity, clinical decision rules, split-sample validation and other tests of validity, as well as economic and decision analyses.14 This level of complexity would almost certainly overwhelm the average ‘target user’ and would preclude most normal mortals from writing guidelines (including one of the present authors); there would also be an issue of time and funding for our volunteers if a complex system of this type were to be introduced. Although these 10 levels of evidence contract down to just four grades of recommendation, feedback from our target audience of colleagues in dermatology mainly suggests that a more ‘user-friendly’ approach is necessary. An excessively rigid and complicated system for making recommendations could, in turn, make guidelines too complex to be of practical value. We are still exploring this issue, looking at the similar grading systems used by organizations such as the Scottish Intercollegiate Guidelines Network15 (based on the U.S. Agency for Health Care Policy and Research guidelines16) and the National Institute for Clinical Excellence (NICE). Even the methods used by these authoritative bodies give rise to some problems as the present NICE publications have some inconsistencies in their evidence grading. For example, their guideline on pressure ulcers17 uses a rather vague three-point scale whereas the guideline on postmyocardial infarction prophylaxis18 has a four-point scale that is more closely linked to defined categories of evidence, although even this allows the lowest (D) grading to be given as an ‘extrapolated recommendation’ from category I evidence (meta-analysis or at least one randomized controlled trial). This year, those in NICE responsible for guidelines will make a policy decision on which evidence grading to utilize (Professor Peter Littlejohns, personal written communication, June 2002), a decision that we will await before making any changes to the BAD guidelines recommendations process, as consistency between guidelines producers would be helpful. It is therefore apparent that the guidelines writing and evaluation process is evolving and progressive, both in overall content as well as regarding the specific information and recommendations of individual guideline topics. The BAD process will continue to look at these wider issues and will be making changes so that its new guidelines will hopefully address developments in validation issues without losing their simplicity. In particular, we will aim routinely to include any financial or conflict of interest information, details of search strategies, a new levels of evidence and grades of recommendation structure, tables or lists of main recommendations, and separate summary sheets. We will reference this article and the earlier description of the overall consultation process8 in future BAD guidelines, as these describe the BAD guidelines consultation process and other relevant background in greater detail than can be included in every single guidelines document.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.066
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Research integrity
Consensus categoriesResearch integrity
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Editorial · Consensus signal: Editorial
Teacher disagreement score0.293
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.066
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.082
GPT teacher head0.424
Teacher spread0.342 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEditorial

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations17
Published2003
Admission routes1
Has abstractyes

Explore more

Same venueBritish Journal of DermatologySame topicClinical practice guidelines implementationFrench-language works237,207