How should we evaluate the risk of bias of physical therapy trials?: a psychometric and meta-epidemiological approach towards developing guidelines for the design, conduct, and reporting of RCTs in Physical Therapy (PT) area: a study protocol
Bibliographic record
Abstract
BACKGROUND: Numerous tools and items have been developed in all health areas to assess the risk of bias of randomized controlled trials (RCTs). The Cochrane Collaboration (CC) released a new tool to assess bias in RCTs, based on empirical evidence quantifying the association between some design features and estimates of treatment effects (TEs). However, this evidence is limited to medicine and investigating a selected set of components. No such studies have been conducted in other health areas such as Physical Therapy (PT) and allied health professions. Evidence specific to the PT area is needed to understand and quantify the association between design features and TE estimates to inform practice and decision-making in this field. The overall goal of this project is to provide direction for the design, conduct, reporting and bias assessment of PT RCTs. We will achieve this through the following specific objectives and methods. METHODS/DESIGN: 1) to measure the association between methodological components and other factors (for example, PT area, type of intervention, type of outcomes) and TE estimates in RCTs in PT, 40 randomly selected meta-analyses of RCTs involving PT interventions will be identified from the Cochrane Database of Systematic Reviews. Trials will be evaluated independently by two reviewers using the most commonly used tools in the PT field. A two-level analysis will be conducted using a meta-meta-analytic approach; 2) to identify relevant items to evaluate risk of bias of PT trials, an exploratory factor analysis (EFA) will be used to identify the latent structure of the items; 3) to develop guidelines for the design, conduct, reporting, and risk of bias assessment of PT RCTs, items obtained from the factor analysis and the meta-epidemiological approach will be further evaluated by experts in PT through a web-based survey following a Delphi procedure. DISCUSSION: The results of this project will have a direct impact on research and practice in PT and are valuable to a number of stakeholders: researchers when designing, conducting, and reporting trials; systematic reviewers and meta-analysts when synthesizing trial results; physiotherapists when making day-to-day treatment decision; and, other healthcare decision-makers, such as those developing policy or practice guidelines.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchMeta-epidemiology (broad) Domain: Methods · Genre: Protocol About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
| gpt | MetaresearchMeta-epidemiology (narrow)Meta-epidemiology (broad) Domain: Methods · Genre: Protocol About the Canadian research system: no · About a Canadian topic: no | Systematic review | medium |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.768 | 0.826 |
| Meta-epidemiology (narrow) | 0.008 | 0.008 |
| Meta-epidemiology (broad) | 0.021 | 0.031 |
| Bibliometrics | 0.022 | 0.017 |
| Science and technology studies | 0.005 | 0.010 |
| Scholarly communication | 0.016 | 0.017 |
| Open science | 0.011 | 0.011 |
| Research integrity | 0.022 | 0.016 |
| Insufficient payload (model declined to judge) | 0.007 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".