A systematic review and development of a classification framework for factors associated with missing patient-reported outcome data
Bibliographic record
Abstract
BACKGROUND/AIMS: Missing patient-reported outcome data can lead to biased results, to loss of power to detect between-treatment differences, and to research waste. Awareness of factors may help researchers reduce missing patient-reported outcome data through study design and trial processes. The aim was to construct a Classification Framework of factors associated with missing patient-reported outcome data in the context of comparative studies. The first step in this process was informed by a systematic review. METHODS: Two databases (MEDLINE and CINAHL) were searched from inception to March 2015 for English articles. Inclusion criteria were (a) relevant to patient-reported outcomes, (b) discussed missing data or compliance in prospective medical studies, and (c) examined predictors or causes of missing data, including reasons identified in actual trial datasets and reported on cover sheets. Two reviewers independently screened titles and abstracts. Discrepancies were discussed with the research team prior to finalizing the list of eligible papers. In completing the systematic review, four particular challenges to synthesizing the extracted information were identified. To address these challenges, operational principles were established by consensus to guide the development of the Classification Framework. RESULTS: A total of 6027 records were screened. In all, 100 papers were eligible and included in the review. Of these, 57% focused on cancer, 23% did not specify disease, and 20% reported for patients with a variety of non-cancer conditions. In total, 40% of the papers offered a descriptive analysis of possible factors associated with missing data, but some papers used other methods. In total, 663 excerpts of text (units), each describing a factor associated with missing patient-reported outcome data, were extracted verbatim. Redundant units were identified and sequestered. Similar units were grouped, and an iterative process of consensus among the investigators was used to reduce these units to a list of factors that met the guiding principles. The list was organized on a framework, using an iterative consensus-based process. The resultant Classification Framework is a summary of the factors associated with missing patient-reported outcome data described in the literature. It consists of 5 components (instrument, participant, centre, staff, and study) and 46 categories, each with one or more sub-categories or examples. CONCLUSION: A systematic review of the literature revealed 46 unique categories of factors associated with missing patient-reported outcome data, organized into 5 main component groups. The Classification Framework may assist researchers to improve the design of new randomized clinical trials and to implement procedures to reduce missing patient-reported outcome data. Further research using the Classification Framework to inform quantitative analyses of missing patient-reported outcome data in existing clinical trials and to inform qualitative inquiry of research staff is planned.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.387 | 0.629 |
| Meta-epidemiology (narrow) | 0.005 | 0.004 |
| Meta-epidemiology (broad) | 0.015 | 0.023 |
| Bibliometrics | 0.079 | 0.052 |
| Science and technology studies | 0.006 | 0.007 |
| Scholarly communication | 0.012 | 0.017 |
| Open science | 0.008 | 0.011 |
| Research integrity | 0.005 | 0.005 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".