Towards Optimum Reporting of Pulmonary Effectiveness Databases and Outcomes (TORPEDO): identifying a core dataset for asthma and COPD studies
Bibliographic record
Abstract
Abstract Purpose There remains a need for a standardized dataset for respiratory studies to accelerate data collection, improve research efficiency and aid the sharing, merging and comparison of datasets. This TORPEDO (Towards Optimum Reporting of Pulmonary Effectiveness Databases and Outcomes) project aimed to develop a checklist of optimum and minimum variables for asthma and chronic obstructive pulmonary disease (COPD) research. Methods A 3-phase modified Delphi survey was conducted: in phase 1, an expert panel generated a list of variables, in phase 2 a Delphi panel selected the minimum variables (>66% agreement) for any design and in phase 3 they were asked to select a minimum set for specific study designs. Results In phase 1 the expert panel (n=22) proposed 224 variables. In phase 2, voting by 64 participants resulted in consensus (>66% agreement) for 18 variables and partial agreement (50-66%) for 44 variables, following this, 5 technical variables (e.g. date of test) were removed. In phase 3, 34 members of the Delphi panel completed voting; consensus was reached for 13 variables for retrospective asthma studies and 34 for prospective asthma studies. For COPD, there were 16 variables for retrospective studies and 37 for prospective studies. Gender, asthma/COPD exacerbations and patient-reported outcomes were the only variables with 100% agreement for both asthma and COPD studies. Conclusion The proposed list of minimally required variables will allow the assessment of current data sources for their utility in asthma and COPD studies, facilitate the merging of datasets, aid standardization of data collection and improve research efficiency.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.719 | 0.704 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.004 | 0.005 |
| Bibliometrics | 0.017 | 0.014 |
| Science and technology studies | 0.004 | 0.005 |
| Scholarly communication | 0.018 | 0.018 |
| Open science | 0.008 | 0.036 |
| Research integrity | 0.003 | 0.006 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".