Bibliographic record
Abstract
We thank Liu et al. for their interest in our paper1 and for initiating an important discussion about using Appraisal of Guidelines for Research and Evaluation (AGREE) tools to evaluate the quality of guidelines. We would like to emphasize that the main purpose of our scoping review was to identify and synthesize outpatient rehabilitation assessment and treatment recommendations for adults continuing to experience signs and symptoms of postacute COVID-19. Secondarily, because we anticipated that several sources of recommendations would not be traditional clinical practice guidelines (given the urgency to provide information to clinicians in a pandemic), we thought it important to provide a relatively simple yet relevant evaluation of the quality and transparency of development of the recommendations identified in the scoping review to assist readers in interpreting our results. However, the main purpose of our study was not to evaluate treatment guideline recommendations. Although we did conduct a critical appraisal of all included studies, we reported results only for the 4 consensus guidelines (Table 3 in our article). All other studies did not meet the majority of our chosen criteria because they were not designed to provide systematically and rigorously developed guidelines. We believe that this result in and of itself (ie, that the majority of recommendations were not developed in a systematic or rigorous manner), provides valuable context for our readers. In determining how best to assess recommendations, we did consider items from the AGREE II as well as from the AGREE-REX. The 2 lead authors completed the AGREE II training tutorials and used the AGREE II in previously published research.2 Although we agree that the AGREE II assesses the quality of the entire guideline development process, its authors state the purpose more broadly as “to provide a framework to: 1) assess the quality of guidelines, 2) provide a methodological strategy for the development of guidelines; and 3) inform what information and how information ought to be reported in guidelines.”3 We found items in the AGREE II that were relevant for our purposes. We did not use the entire AGREE II tool because we anticipated that many of the items would not be relevant for the recommendation papers we were likely to find. We acknowledge that it would be more accurate to state that we used selected items from the AGREE II to inform our critical appraisal. It was not our intention to imply that we used the AGREE II tool in its entirety or that we followed recommended practices (eg, rating each statement on the designated 7-point scale). Rather, we thought that the answers to the standardized items (selected from 3 of the 6 AGREE II domains) would help clinicians to better understand the background of recommendations provided in the identified literature. This is why we did not report a numerical score, but rather yes/no answers to the statements selected from the AGREE IItool. In their letter, Liu et al. have suggested that the AGREE II focuses on methodological quality and does not evaluate the evidence behind the recommendations. In assessing the quality of guidelines (the first stated purpose of AGREE II),2 the tool does provide some statements to evaluate evidence (in Domain 3, Rigour of Development). For example, item #9 asks the assessor to determine whether “the strengths and limitation of the body of evidence are clearly described”; item #12 asks whether “there is an explicit link between the recommendations and the supporting evidence”; and item #13 asks whether “the guideline has been externally reviewed by experts prior to its publication”.2 We included statements #9 and #13 in the 7 questions we selected to use in our critical appraisal. Liu et al. have suggested that it may have been more appropriate for us to use the AGREE-REX in our study. We did review the AGREE-REX before deciding on our methods. Both the AGREE II and the AGREE-REX deal with evaluating quality of guidelines,3,4 and the authors of the AGREE-REX state that the tool is meant to complement the AGREE II by including items that address clinical credibility, consideration of values of all stakeholders, and implementability of the recommendations.3 In evaluating the criteria listed for each item in the AGREE-REX, we found some overlap with items from the AGREE II. For example, AGREE-REX #1 provides 8 criteria that can be used to evaluate the evidence supporting the recommendations. These criteria are similar to criteria included in items #7 and #9 from the AGREE II. As another example, AGREE-REX #5 includes 4 criteria related to the values and preferences of patients/populations. These criteria are similar to AGREE II #5, which asks whether the views and preferences of the target population have been sought. Overall, we found quite a bit of similarity between items that we thought were relevant for our purposes from the AGREE II and AGREE-REX tools. We felt the wording of the items chosen from the AGREE II was straightforward and provided enough clarity for clinicians to understand aspects related to the quality and transparency of development of the recommendations. We thank Liu et al. for drawing attention to the different AGREE tools available to evaluate guidelines relevant to clinical practice. We agree that it would not be appropriate to suggest that using only select questions from the AGREE II—and providing only a yes/no evaluation (vs rating on the recommended 7-point scale after assessing all suggested criteria for each item)—could provide the same degree of information about guideline quality and development as that obtained through proper use of the entire tool. Our intention was simply to conduct a brief appraisal to provide some context for clinicians to better understand the included studies, according to a simplified set of criteria that we did not develop on our own but that we extracted from items in the AGREE II tool.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.113 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.005 | 0.004 |
| Scholarly communication | 0.009 | 0.003 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.059 | 0.033 |
| Insufficient payload (model declined to judge) | 0.059 | 0.046 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".