Quality Criteria for Real-world Data in Pharmaceutical Research and Health Care Decision-making: Austrian Expert Consensus
Bibliographic record
Abstract
Real-world data (RWD) collected in routine health care processes and transformed to real-world evidence have become increasingly interesting within the research and medical communities to enhance medical research and support regulatory decision-making. Despite numerous European initiatives, there is still no cross-border consensus or guideline determining which qualities RWD must meet in order to be acceptable for decision-making within regulatory or routine clinical decision support. In the absence of guidelines defining the quality standards for RWD, an overview and first recommendations for quality criteria for RWD in pharmaceutical research and health care decision-making is needed in Austria. An Austrian multistakeholder expert group led by Gesellschaft für Pharmazeutische Medizin (Austrian Society for Pharmaceutical Medicine) met regularly; reviewed and discussed guidelines, frameworks, use cases, or viewpoints; and agreed unanimously on a set of quality criteria for RWD. This consensus statement was derived from the quality criteria for RWD to be used more effectively for medical research purposes beyond the registry-based studies discussed in the European Medicines Agency guideline for registry-based studies. This paper summarizes the recommendations for the quality criteria of RWD, which represents a minimum set of requirements. In order to future-proof registry-based studies, RWD should follow high-quality standards and be subjected to the quality assurance measures needed to underpin data quality. Furthermore, specific RWD quality aspects for individual use cases (eg, medical or pharmacoeconomic research), market authorization processes, or postmarket authorization phases have yet to be elaborated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.069 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".