Bibliographic record
Abstract
This issue of The Canadian Journal of Psychiatry (The CJP) includes a Perspective article written by Dr H Edmund Pigott.1 Dr Pigott articulates a series of criticisms of the Sequenced Treatment Alternatives to Relieve Depression (STAR*D) trial. Some of these arguments attack the trial methodology and reporting, whereas others take issue with the broader concept of measurement-based care. There is a reason for the ascendancy of RCTs. The process of randomization helps to ensure that potential confounding variables are equally distributed between treatment groups. Unique among procedures for controlling confounding (for example, matching and statistical modelling), randomization controls for unmeasured, and even unknown, confounders. As such, it is a uniquely powerful study design. The CJP’s Perspectives format is intended to be open to expressions of opinion. Authors of Perspectives articles are encouraged to take positions on controversial topics,2,3 as Dr Pigott has done. In an accompanying Guest Editorial,4 Dr Raymond W Lam and Dr Sidney H Kennedy take issue with 3 of the arguments put forward by Dr Pigott: the assertion that a primary objective of the STAR*D trial was modified post hoc; the idea that remission is not a suitable outcome measure in depression research; and the idea that the STAR*D story provides a general indictment of the principles of measurement-based care. Questions surrounding modification of a priori hypotheses in clinical trials are not trivial. Statistical analyses are vulnerable to error if there is a lack of clarity surrounding their primary outcomes. Problems also arise when primary outcomes are not clearly distinguished from secondary and exploratory ones. In the worst-case scenario, data are analyzed and investigators then selectively report the findings that they prefer. This is the proverbial fishing expedition seen in undisciplined research. In this disastrous scenario, readers cannot correctly interpret reported P levels or confidence intervals as these no longer reflect their intended probabilities. Instead, they partially reflect decisions made by the authors. This does not mean that investigators should be denied full rein to explore their data. However, if exploratory analyses are to be conducted, they must be carefully identified as such. Results from exploratory analyses are provisional. They require additional replication to ensure that the results were not merely statistical outliers. In their Guest Editorial, Lam and Kennedy4 argue that Pigott has misunderstood some aspects of the STAR*D protocol. They feel that outcomes in question, those for the Quick Inventory of Depression Symptoms—Self-Report remission, were adequately characterized as post hoc analyses, whereas Pigott feels that these were obscured. This issue is one of transparency of reporting. Progress is being made in addressing such problems. For example, The CJP requires that clinical trials be registered in a suitable archive at or before the onset of subject enrolment. A suitable registry must be accessible to the public at no charge; be open to all prospective registrants; be managed by a not-for-profit organization; have a mechanism to ensure the validity of the registration data; and be electronically searchable. Registration protects the integrity of trials, both from methodological manipulation and from accusations of such manipulation. However, registration, in itself, is not always enough. After all, the STAR*D trial protocol was registered,5 but this did not prevent concerns and controversies from arising. Some journals (for example, The Journal of the American Medical Association) now require clinical trial protocols, including the complete statistical analysis plan, be submitted along with submissions of clinical trial reports. Increasingly, investigators are choosing to archive their complete RCT protocols, including their detailed a priori analysis plans. In contrast, Dr Pigott reports that he needed a Freedom of Information Act request to obtain the STAR*D protocol. In the systematic review literature, it is a long-standing practice to publish or archive review protocols prior to the reviews being conducted. The CJP’s recently initiated Systematic Reviews category strongly encourages authors to register their review protocols in a suitable registry (for example, PROSPERO) or to archive or publish their full protocols.3 While most major journals, this one included, require registration for trials, they usually do not require it for other types of studies. Ultimately, it would be a good idea for authors of all studies that use statistical analysis, such as epidemiologic studies or brain imaging studies, to register their protocols in a similar fashion. This would assist readers in their interpretation of the statistics reported in the ensuing papers. It would also protect them from later accusations that they have modified their analysis after the fact. While they disagree on many points, the authors of these papers1,4 agree that the STAR*D results were disappointing. Other papers in this issue of The CJP amplify this sense of urgency for greater progress against depression. Sakina J Rizvi and collaborators6 report unemployment–disability rates of 30.3% of a sample of depressed primary care patients and of 41.4% in a tertiary care sample. An updated description of the general population’s prevalence of major depressive disorder documents continuing high prevalence,7 associated dysfunction, comorbidity, and stigmatization.8 Better publication and reporting standards will help us to keep our sights on the true enemy: depression itself.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.036 | 0.076 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.013 | 0.028 |
| Scholarly communication | 0.022 | 0.035 |
| Open science | 0.004 | 0.013 |
| Research integrity | 0.028 | 0.057 |
| Insufficient payload (model declined to judge) | 0.029 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".