MétaCan
Menu
Back to cohort
Record W4383265828 · doi:10.1186/s12916-023-02937-0

Development of consensus-driven SPIRIT and CONSORT extensions for early phase dose-finding trials: the DEFINE study

2023· review· en· W4383265828 on OpenAlexafffund
Olga Solovyeva, Munyaradzi Dimairo, Christopher J. Weir, Siew Wan Hee, Aude Espinasse, Moreno Ursino, Dhrusti Patel, Andrew Kightley, Sarah Hughes, Thomas Jaki, Adrian Mander, T.R. Jeffry Evans, Shing Yip Lee, Sally Hopewell, Khadija Rantell, An‐Wen Chan, Alun Bedding, Richard Stephens, Dawn P. Richards, Lesley Roberts, John P. Kirkpatrick, Johann S. de Bono, Christina Yap

Bibliographic record

VenueBMC Medicine · 2023
Typereview
Languageen
FieldMedicine
TopicEthics in Clinical Research
Canadian institutionsRobarts Clinical TrialsCanada Research ChairsWomen's College HospitalUniversity of Toronto
FundersGenentechNational Institutes of HealthSierra OncologyAstellas PharmaEisaiNational Center for Advancing Translational SciencesMedical Research CouncilDepartment of Health and Social CareNational Institute for Health and Care ResearchCancer Research UKPlexxikonSeagenPfizerNuCanaSanofiCelgenePTC TherapeuticsUK Research and InnovationHalozymeMedical Research Charities GroupBeiGeneGlaxoSmithKlineEli Lilly and CompanyBristol-Myers Squibb
KeywordsConsolidated Standards of Reporting TrialsMedicineDelphi methodDelphiClinical trialProtocol (science)Thematic analysisAlternative medicineMedical educationFamily medicineQualitative researchComputer sciencePathologySocial scienceArtificial intelligence

Abstract

fetched live from OpenAlex

BACKGROUND: Early phase dose-finding (EPDF) trials are crucial for the development of a new intervention and influence whether it should be investigated in further trials. Guidance exists for clinical trial protocols and completed trial reports in the SPIRIT and CONSORT guidelines, respectively. However, both guidelines and their extensions do not adequately address the characteristics of EPDF trials. Building on the SPIRIT and CONSORT checklists, the DEFINE study aims to develop international consensus-driven guidelines for EPDF trial protocols (SPIRIT-DEFINE) and reports (CONSORT-DEFINE). METHODS: The initial generation of candidate items was informed by reviewing published EPDF trial reports. The early draft items were refined further through a review of the published and grey literature, analysis of real-world examples, citation and reference searches, and expert recommendations, followed by a two-round modified Delphi process. Patient and public involvement and engagement (PPIE) was pursued concurrently with the quantitative and thematic analysis of Delphi participants' feedback. RESULTS: The Delphi survey included 79 new or modified SPIRIT-DEFINE (n = 36) and CONSORT-DEFINE (n = 43) extension candidate items. In Round One, 206 interdisciplinary stakeholders from 24 countries voted and 151 stakeholders voted in Round Two. Following Round One feedback, one item for CONSORT-DEFINE was added in Round Two. Of the 80 items, 60 met the threshold for inclusion (≥ 70% of respondents voted critical: 26 SPIRIT-DEFINE, 34 CONSORT-DEFINE), with the remaining 20 items to be further discussed at the consensus meeting. The parallel PPIE work resulted in the development of an EPDF lay summary toolkit consisting of a template with guidance notes and an exemplar. CONCLUSIONS: By detailing the development journey of the DEFINE study and the decisions undertaken, we envision that this will enhance understanding and help researchers in the development of future guidelines. The SPIRIT-DEFINE and CONSORT-DEFINE guidelines will allow investigators to effectively address essential items that should be present in EPDF trial protocols and reports, thereby promoting transparency, comprehensiveness, and reproducibility. TRIAL REGISTRATION: SPIRIT-DEFINE and CONSORT-DEFINE are registered with the EQUATOR Network ( https://www.equator-network.org/ ).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Direct model labels (unvalidated)

Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.

Model armCategoriesStudy designConfidence
gemmaMetaresearch
Domain: Reporting · Genre: Methods
About the Canadian research system: no · About a Canadian topic: no
Not applicablelow
gptMetaresearch
Domain: Reporting · Genre: Methods
About the Canadian research system: no · About a Canadian topic: no
Other designhigh
models splitAgreement compares identical category sets and study designs across arms.

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.043
metaresearch head score (Gemma)0.212
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow)
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.846
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0430.212
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0060.000
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.951
GPT teacher head0.729
Teacher spread0.223 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Labeled directly by 2 models reading the full record.

Metaresearch

The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.

Study designNot applicable · Other design
DomainReporting
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations12
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueBMC MedicineSame topicEthics in Clinical ResearchCategoryMetaresearchFrench-language works237,207