MétaCan
Menu
Retour à la cohorte
Enregistrement W1997885956 · doi:10.1097/ta.0b013e318256dc4d

Evidence level of individual studies

2012· review· en· W1997885956 sur OpenAlexaboutno aff
Angela Sauaia, Ernest E. Moore, Jennifer Crebs, Ronald V. Maier, David B. Hoyt, Steven R. Shackford

Notice bibliographique

RevueThe Journal of Trauma: Injury, Infection, and Critical Care · 2012
Typereview
Langueen
DomaineDecision Sciences
ThématiqueMeta-analysis and systematic reviews
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésPsychology

Résumé

récupéré en direct d'OpenAlex

Evidence-based medicine is “the conscientious, explicit, and judicious use of current best evidence in making decisions about the care of individual patients.”1 This becomes complicated when busy health care providers are faced with the task of summarizing “the current best evidence.” Systematic reviews, such as those published by the Cochrane Collaboration,2 serve this purpose by appraising and distilling the daunting amount of information available. These reviews are commonly attached to a level of evidence, which gauges the confidence of estimates reported by existing studies. Thus, levels of evidence pertain to the knowledge generated by the summative collection of research on a specific topic. Although the evidence base can come from a single study, more often it is the final step of a long scientific journey, in which experts collect, appraise, and summarize the findings of several individual studies using a specific, standard methodology.1 Different systems to define hierarchy of evidence have been proposed by renowned groups including the pioneer Cochrane Collaboration,2 the Oxford Centre for Evidence-Based Medicine (OCEBM),3 the US Preventive Task Force,4 and the Evidence-Based Practice Center (EPC) program of the US Agency for Healthcare Research and Quality.5 The recently launched Grades of Recommendation, Assessment, Development, and Evaluation (GRADE)6 system follows a detailed stepwise process to rate evidence and to determine the strength of recommendations in systematic reviews, health technology assessments, and clinical practice guidelines. According to the GRADE group Web site, more than 50 organizations have endorsed their system, including the World Health Organization, the American College of Physicians, the American College of Chest Physicians, the American Endocrine Society, the American Thoracic Society, the Canadian Agency for Drugs and Technology in Health, and the UK’s National Institute for Health and Clinical Excellence. The British Medical Journal encourages authors of clinical guidelines to use the GRADE system.7 The Cochrane Collaboration has also adopted the principles of the GRADE system for evaluating the quality of evidence for outcomes reported in systematic reviews.8 Yet, a systematic review is not always available; thus, how will busy health care providers manage the formidable volume of information that becomes available everyday? “How does the article I read today change (or not) what I will recommend to my patients tomorrow?” To address this imperative, several scientific journals have recently adapted grading systems to assess the level of evidence of individual articles in an effort to provide guidance to their readers.9–12 This year, The Journal of Trauma joined the discourse by requiring authors to assign levels of evidence to their own clinically oriented studies. As detailed previously, the existing grading systems (e.g., GRADE) were originally designed to rate the summative body of evidence and not the level of evidence in individual articles. The grading of evidence in individual studies is a middle step in determining the hierarchy of evidence and comprises a judgment regarding the confidence and uncertainty emanating from a particular study. Several aspects of the investigation are examined, including, but not limited to appropriate design to address well-formulated research questions, appropriately measured outcomes, assessment of inferential error, risk of bias, and control of confounding. The results of this appraisal will inform the reader about the level of uncertainty of the study’s findings and how much it adds to the existing knowledge in that topic. As part of the process of assessing the overall level of evidence, the GRADE system rates the evidence from individual studies into one of four categories ranging from high to very low (Table 1).13 Study design is the GRADE’s critical measure to classify the quality of the evidence: for therapeutic studies, randomized clinical trials (RCTs) always start as High and observational studies as Low. From this starting point, evidence may be downgraded or upgraded through the evaluation of several specific domains as follows: (1) risk of bias, (2) imprecision, (3) inconsistency, (4) indirectness, (5) publication bias, (6) effect size, (7) existence of dose-response pattern, and (8) effect of plausible confounding on findings (Table 2). When the study addresses diagnostic accuracy, however, a slightly different classification applies.14TABLE 1: GRADE Quality Assessment CriteriaTABLE 2: Factors That May Decrease or Increase the Quality of EvidenceGRADE is a somewhat complicated system, which, as its own authors recognize, involves an element of subjectivity.15 In addition, there is the negative connotation created by classifying a study as low quality, a term which could be construed as lack of scientific rigor on the part of the authors. As the PRISMA authors wisely put it, “quality is often the best the authors have been able to do,” and recommended the term risk of bias instead.16 Quality should be assessed when accepting or rejecting a manuscript for publication, whereas the evidence level of individual studies involves judging a study’s level of uncertainty and risk of bias. Rather than introducing a new term (such as risk of bias), with which readers and authors may not be familiar, we propose to use the established nomenclature “evidence level of individual studies” (ELIS). Consensus statements such as the GRADE system and the OCEBM guidelines can serve as the basis upon which to build a standard to assess ELIS. The proposed ELIS system retains study design as a major factor in the classification but recognizes that each type of clinical question (therapeutic, diagnostic accuracy, etc.) demands different types of study designs. The proposed ELIS framework (Table 3) is heavily based on the previous groundbreaking work from the GRADE workgroup, the OCEBM 2009 and 2011 guidelines, and the Journal of Bone and Joint Surgery’s adaptation of OCEBM’s materials, which have been well accepted by the scientific community and shown to have acceptable reliability.17,18 The determination of ELIS involves three steps.TABLE 3: Proposed Evidence Level of Individual Studies (ELIS)ELIS STEP 1: DEFINE STUDY TYPE Therapeutic and care management studies evaluate a treatment efficacy, effectiveness, and/or potential harm, including comparative effectiveness research and investigations focusing on adherence to standard protocols, recommendations, guidelines, and/or algorithms. Prognostic and epidemiologic19 studies assess the influence of selected predictive variables or risk factors on the outcome of a condition. These predictors are not under the control of the investigator(s). Epidemiologic investigations describe the incidence or prevalence of disease or other clinical phenomena, risk factors, diagnosis, prognosis or prediction of specific clinical outcomes, and investigations on the quality of health care. Diagnostic tests or criteria 20 studies describe the validity and applicability of diagnostic tests/procedures or of sets of diagnostic criteria used to define certain conditions (e.g., definition of adult respiratory distress syndrome, multiple organ failure, or postinjury coagulopathy). Economic and value-based evaluations focus on which type of care management can provide the highest quality or greatest benefit for the least cost. Several types of economic evaluation studies exist, including cost-benefit, cost-effectiveness, and cost-utility analyses. More recently, Porter21,22 proposed value-based health care evaluations, in which value was defined as the health outcomes achieved per dollar spent. Systematic reviews and meta-analyses (SR/MA) evaluate the body of evidence on a topic; meta-analyses specifically include the quantitative pooling of data. Guidelines are systematically developed statements to assist practitioner and patient decisions about appropriate health care for specific clinical circumstances.23 ELIS STEP 2: DEFINE THE RESEARCH DESIGN Table 3 reflects a different hierarchy of designs for each of the previously mentioned study types. For therapeutic studies, RCTs remain the paragon of biomedical research, and other treatment study designs will still rank lower than level I evidence. This is because the processes used to conduct RCTs minimize the risk of confounding factors influencing the results. As a result, the findings generated by RCTs are likely to be closer to the true effect than the findings generated by other research methods. Prognostic studies allow for more flexibility in study design, with cohort prospective studies with preestablished hypotheses generating less uncertainty and consequently stronger evidence than those of case-control designs. It is important that we clearly define case-series versus comparative, cohort versus case-control, and prospective versus retrospective studies. Case-series studies evaluate a group of patients submitted to a type of care/procedure/test without a suitable comparison group. A comparison group can be a group of patients with similar characteristics who received a different type of care/procedure/test or, alternatively, the investigator can compare the same group of patients before and after an intervention. It is not difficult to realize that the lack of a comparator makes us less confident in the evidence and less likely to adopt the new procedure. Of course, if this is an innovative treatment of a lethal disease for which there is no available treatment, we may adopt it even with low confidence for lack of better options. The urgency to adopt the new treatment, however, does not change the fact that our confidence is still low and that further research will be crucial to increase our confidence level. Once we determine that there is a comparator group, we can define whether this is a cohort or case-control study. The fundamental difference between these two designs lies on when the investigators determine the exposure/risk factor and the outcome.3 In case-control studies, the outcome is determined first and the exposure/risk factor/intervention later. For example, Wu et al.24 used a case-control design to compare the bone mineral density of 87 elderly patients with hip fractures to 87 elderly patients without hip fractures and found it to be significantly lower in the first group. In cohort studies, investigators define first the exposure/risk factor and then assess their outcome of interest. For example, Lin et al. enrolled a cohort of 217 elderly patients with hip fractures in their study and evaluated a risk factor defined as body mass index ratio between the greater trochanter and the femoral neck. In case-control studies, a group with the outcome (the “cases”) is compared with a group without the outcome but otherwise similar (the “controls”) regarding something that happened to them before they experienced the outcome. In cohort studies, a group of patients with at least one common characteristic (the “cohort”) is assessed for the development of outcome(s). This distinction can get confusing when the authors compare, for example, survivors to nonsurvivors regarding a specific risk factor. In this case, readers will know that the study is a cohort if both survivors and nonsurvivors were consecutive patients with a common risk factor (e.g., trauma). To make things more confusing, a case-control study is sometimes a later offspring of a well-planned cohort study, as in the case-control study by Shaz et al.25 on postinjury coagulopathy. Although some of the most important medical discoveries were done through case-control studies,26 this design has several limitations that place it lower on the evidence hierarchy. These include, but are not limited to, potential for bias in the selection of the control group and uncontrolled confounding in the assessment of the risk factor/exposure. In this proposed ELIS, the terms prospective and retrospective refer to the intention underlying data collection, rather than when data were actually retrieved. If the data were compiled to answer a predefined set of research questions, then this is a prospective study, regardless of whether data were accrued concurrently with care or after the fact through records review. Conversely, the use of data to answer a question unrelated to the original question for which the data were gathered is a retrospective analysis. Some clinical databases are prospectively planned to answer a broad set of predetermined questions (e.g., the Denver MOF database constructed to assess early risk factors for postinjury multiple organ failure).27 Information recorded for other purposes (e.g., medical records, operating room registries, claims data) can only produce a retrospective analysis. Disease registries (e.g., trauma registries) are a point of contention because data collection occurs concurrently with care (commonly reported as “data were prospectively collected” or “patients were prospectively included into a registry”). The major strength of these registries lies on the quality of the data collected because factors such as recall bias and missing data are less likely. Yet, they can generate both prospective studies of preestablished outcomes and retrospective studies when used for not preestablished outcomes. Why is this distinction (retrospective vs. prospective) important in establishing the ELIS? It is important because retrospective studies are more subject to biases (e.g., relevant variables may not have been included or were measured using different methods) that decrease our confidence in their results. Furthermore, retrospective multiple unplanned comparisons increase the potential of a type I error, as explained in greater detail later in this article. ELIS STEP 3: ASSESS THE STRENGTHS AND LIMITATIONS OF THE STUDY THAT WILL AFFECT THE UNCERTAINTY OF THE RESULTS The next step in determining the ELIS recognizes that all research designs, even RCTs, are more or less limited by confounding, bias, inadequate sample size and statistical power, heterogeneity of included populations, differences between control and study groups, missing data, loss to follow up, and so on. All these factors affect the uncertainty around study outcomes. We combined some of the GRADE-defined factors (Table 2) with the earlier OCEBM table (Table 4) to modify the ELIS.TABLE 4: Oxford Center for Evidence-Based Medicine Levels of Evidence (March 2009)To define the magnitude of effect, we assessed the size of the relative risk (RR) within the context of disease severity. Thus, for a moderately severe condition (with low-to-moderate morbidity/mortality), a large effect was defined as a high RR (>5 or <0.2), whereas for more severe diseases, only a moderate-to-large RR (2–5 or 0.2–0.5) was required. The statistical power of the study is a critical aspect in determining ELIS. It is usually easier to first define the situations where statistical power is not important: once the study detects a significant difference for an a priori stated hypothesis, the issue of statistical power is irrelevant. When investigators conduct multiple unplanned comparisons, then the potential for type I error (the error of finding a difference when in fact there is not one, usually set at <0.05) increases. Statistical power becomes relevant when no significant difference is detected and we want to gauge the type II error (the error of not finding a difference when in fact there is one). When this happens, authors may declare “failure to detect a significant difference” and should provide the statistical power for detecting the observed (or predetermined) difference (generally accepted as adequate when >80%). This is often the case when assessing whether randomization was successful and the two RCT groups do not show statistical differences; or for secondary outcomes for which the study was not powered. In an alternative scenario, which is becoming more common with the popularity of comparative effectiveness studies, the authors may aim at declaring bioequivalence or noninferiority. In this case, power must be determined with as much rigor as we usually determine significance, thus requiring levels greater than 90%. Of course, akin to the well-known p < 0.05, statistical power levels are arbitrary and should reflect the specific topic of the study. Studies of lethal conditions without a known treatment may require lower confidence levels (e.g., p < 0.10 or p < 0.15) to establish a significant difference or lower statistical power to declare noninferiority whereas investigations of low morbidity/mortality conditions with established treatments may require higher power or confidence levels. Furthermore, differences can be statistically significant and clinically meaningless. In sum, statistical power and confidence are functions of the clinical question being answered by the study, not after-the-fact considerations. Other ELIS modifiers were included in two sets of “negative criteria” at the bottom of Table 3. One set is for general types of studies and includes confounding, bias, loss to follow-up, missing data, and heterogeneity of the populations. We contemplated using dose-response as factor, as recommended by the GRADE group, but decided that this was a difficult element to define in the instructions for authors. Especially in trauma and acute care, assessment of dose-response patterns can be complicated by survivorship bias.28 Instead, we encourage our reviewers to take dose-response gradients into consideration on a case-by-case basis. As mentioned previously, heterogeneity of populations must be taken into consideration when appraising a study, particularly multi-institutional studies (even when RCT is the design), studies including condition(s) caused by different pathogenic mechanisms (e.g., patients with sepsis, patients with critical illness), and national and international disease is also a major in thus, this element was to the set of negative specific to A final to to the quality, and validity of collected data. Especially when studies with large data sets recorded by multiple we encourage authors to these (e.g., of the records were and was assessed by the For diagnostic studies, the use of a standard is the factor. When all patients with a condition are submitted both to the (or set of diagnostic under investigation and the the is a when only a group of patients with the condition (e.g., patients who are more the is submitted to the standard The quality of the standard of course, a We are all there are no of a between outcomes. In addition, the of the standard is sometimes (e.g., all patients with to and (e.g., for all Yet, this is a that our confidence in the thus, it must be in the ELIS. ELIS AND The ELIS can only be assessed if all are included in the The article must all information for the study to be by including and randomization the case of confounding risk potential for bias, and statistical analysis. For that guidelines etc.) is The the Quality and Of health Research Web is an of and we recommend it to authors articles to the Journal of This ELIS classification does not reflect the scientific rigor or research of the study. A study that is because it not follow scientific does not new and the of publication should be very ELIS, we propose to gauge some of the uncertainty of a study’s which will its to current In addition, the ELIS must be used in with the determination of that whether the patients and outcomes in the are similar to of the results to the patient It is that an not to define study as evidence.” level II or evidence, however, should not be as in health and health care, such as and to the of Study designs reflect the of and In a on the limitations of RCTs in the care and the of other study designs in the care Yet, these does not change the level of uncertainty with specific study designs, limited risk and high risk of bias. Thus, our proposed ELIS system does not to define the of a study but rather to its level of In addition, we propose a new framework for the in which the authors place their results in context and the ELIS of their study. We authors to use the to assist health care providers in the question stated at the of this “How does the article I read today (or not) what I will recommend to my patients tomorrow?” We encourage authors to describe how the study to existing knowledge about the topic and provide guidance on how results should be used by their readers in their current clinical practice using the As investigators should also propose new studies likely to increase the level of evidence of existing In sum, we have proposed a system to the level of uncertainty of individual studies to the of studies, those with care. We to from our The authors declare no of interest.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Étiquettes directes de modèles (non validées)

Étiquettes de catégorie et de devis d'étude par modèle, issues des rondes d'étiquetage. C'est une sortie machine, non validée, et le désaccord entre modèles est livré comme donnée. Aucun devis ici n'est encore validé contre MEDLINE.

BrasCatégoriesDevis d'étudeConfiance
gemmaMétarecherche
Domaine: Évaluation · Genre: Synthèse
Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non
Théorique ou conceptuellow
gptMétarecherche
Domaine: Évaluation · Genre: Synthèse
Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non
Théorique ou conceptuelmedium
modèles en accordL'accord compare des ensembles de catégories et des devis identiques entre les bras.

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,094
score de la tête « metaresearch » (Gemma)0,053
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche, Charge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesMétarecherche
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Autre devis · Signal consensuel: aucune
GenreSignal candidat: Synthèse · Signal consensuel: Synthèse
Score de désaccord entre enseignants0,971
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0940,053
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0080,003
Bibliométrie0,0010,001
Études des sciences et des technologies0,0000,001
Communication savante0,0000,001
Science ouverte0,0020,000
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0010,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,947
Tête enseignante GPT0,646
Écart entre enseignants0,301 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Étiqueté directement par 2 modèles lisant le dossier complet.

Devis d'étudeThéorique ou conceptuel
DomaineÉvaluation
GenreSynthèse

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations28
Publié2012
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueThe Journal of Trauma: Injury, Infection, and Critical CareMême sujetMeta-analysis and systematic reviewsCatégorieMétarechercheTravaux en français237 207