Next generation health guidelines: The role of real‐life data in evidence‐based medicine
Bibliographic record
Abstract
Adaptation processes are used to improve the efficiency and applicability of health guidelines. Adaptation using source guidelines developed by others to create context-specific recommendations are of increasing interest. Real-life data (RLD) such as data from registries, electronic health records, or patient-generated data play an important role in this process and facilitate the implementation of the guidelines. However, there is a need to define RLD, its generation and integration in the evidence to decision processes. This editorial aims to inform the scientific community on the initiative set forth by the European Academy of Allergy and Clinical Immunology (EAACI)- Respiratory Effectiveness Group (REG) on March 15, 2023, in Lisbon. The growing interest in “Real World Evidence” is undeniable in several allergic and respiratory diseases. REG and EAACI, within the Presidential initiative EAACI Real Life (EARL) and through its Methodology and Research and Outreach Committees, are committed to the responsible use of data that reflect the real-life of patients. This editorial reports on the workshop held by the REG and EAACI on March 15, 2023, and intents to clarify the role of what is widely described of real-world evidence (RWE) and how it can support the development and implementation of health guidelines.1, 2 Well-planned and executed randomized controlled trials (RCT) are the cornerstone of evidence-based medicine1, 2 and are used by regulatory authorities, such as the Food and Drug Administration (FDA) and European Medicines Agency (EMA) to make informed decisions. There is concern within the scientific and practice community that RCTs do not reflect the general population as they are performed on selected groups under controlled settings. To overcome this limitation, non-randomized studies of interventions (NRSI) may be used as a source of complementary, sequential or replacement evidence3 for RCTs. Together, or individually, these studies generate a body of evidence that should mimic real-life situations. Thus, we endorse RLD to support decisions that reflect real-life scenarios. The REG “Manifesto on the importance of real-life research”4 endorsed by other scientific organizations, was a first initiative describing RLD. A new tool—Real Life Evidence AssesmeNt Tool (RELEVANT) was developed in collaboration with EAACI to evaluate real-world studies. However, it is important to clarify that the RWE concept is a misconception as all data belongs to the real world, regardless of if they are generated by RCTs or NRSI1, 2 since the FDA considers any study design as being able to provide real-life data. The limitations in study design, execution, and applicability determines how studies inform the decision making in a given context. The term real-world can be considered synonymous with the attempt to optimize applicability, provided that there is an understanding that it may increase the risk of bias (ROB), particularly if studies are not randomized. Taking these points into account EAACI and REG held a workshop on the optimization of RLD inclusion in support of the recommendations included in health guidelines.3 The following organizations participated: Allergy Scientific Societies (American Academy of Allergy, Asthma and Immunology, World Allergy Organization, Asian Pacific Association of Allergy, Asthma and Clinical Immunology, Latin American Society of Allergy and Immunology), Respiratory Societies (American Thoracic Society, American College Chest Physicians, Interasma, International Primary Respiratory Group), Guidelines groups (Allergic Rhinitis and its Impact on Asthma, Global Initiative in Asthma, Global Initiative for Chronic Obstructive Lung Disease, Guidelines International Network), regulatory authorities (FDA, EMA and Paul Erlich Institute, National institute for Health and Care Excellence) in addition to Observational and Pragmatic Research Institute and the patients' organization Global Allergy & Asthma Patients. The following sessions were held to address the challenges: In the first session RCT-derived evidence contribution to RLD was extensively analyzed. Key topics included critical appraisal of the current guidelines adaptation process, alignment of the “certainty level” of RLD, and the incorporation of NRSI into evidence-to-decision frameworks.5 The second session focused on mitigating the ROB in the context of RCTs or NRSI. The GRADE approach evaluates the ROB by addressing limitations in study design and execution (internal validity), external validity (or indirectness), inconsistency of results, and publication bias. This session also emphasized the importance of registration of all study protocols, including NRSI, and proposed it as a mandatory step to reach data integrity. In the third session the growing role of registries was recognized.6-9 The importance of properly designed registry systems was highlighted as key for their validity. Severe asthma is a good example of the continuous scientific and clinical inputs provided by national and international registries. Nonetheless, there is a pressing need in developing a correct methodology for creating and structuring registries and for achieving sustainability by facilitating data collection in a busy clinical practice. Additionally, it was acknowledged that registries can apply randomization to increase validity. A general agreement was reached on the importance to critically evaluate the RLD collected through mobile apps. A consensus was attained on the relevance and standards for RLD and their integration into health guidelines. When addressing data-quality issues, there are several dimensions to consider, including the depth, breadth, coverage, timeliness, and potential impact of the data. These can be incorporated in an automated analysis framework evaluating their reliability, usability, and compliance. Central data-procurement teams may help resourcing efficiency, eliminate silos, and prevent duplication. The future role of artificial intelligence to assess the quality of RLD was acknowledged. Our next goal is to develop general and shared rules on the approaches. The next workshop is scheduled on October 4 in Rome and Allergy readers will be updated promptly about the topics discussed and future directions. All authors equally contributed to the writing and critical revision of the editorial. None. There are no conflicts of interest related to this manuscript. Data sharing is not applicable to this article as no new data were created or analyzed in this study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".