Does network meta-analysis generate any new knowledge?
Bibliographic record
Abstract
doi:10.1160/TH12-10-0749 Thromb Haemost 2013; 109: ■■■■ Dear Sirs, To better understand the point of view presented by Schulman in his recent editorial (1), three points may deserve further comment and clarification. The first is whether network meta-analysis (NETMA) studies should be seen as a big advancement in the knowledge about a clinical problem or as a small one. Schulman is absolutely correct when he points out that the incremental knowledge generated by NETMA is small in comparison with an analysis based on purely narrative approaches or on rudimental, but effective algebraic calculations. For this reason, our approach towards NETMA (2-4) is to present our results exclusively in very short papers (mainly, Letters to the Editors of less than 500 words); in fact, full-length papers of 1,500 words or more would not be justified by the limited advancement in the understanding of the problem. The second point regards the simplified calculation presented by Schulman in his Table 3 (1). Typically, all NETMA studies generate a point estimate for the concerned parameter (e.g. the odds ratio [OR] for indirect comparisons) along with a measure of the statistical variability of the point estimate (e.g. the 95% confidence interval of the OR and/or the p-value). While, on the one hand, there is a lot of mathematical complexity behind the calculations of statistical variability, on the other hand point estimates are devoid of any complexity at least in the simplest models of NETMA. As a result, when Schulman proposes his simple method of calculation (see Table 3 in [1]), actually he is already using the basic formula of NETMA as implemented in the Bucher method (5) and in the ITC software (Canadian Agency for Drugs and Technologies in Health, Indirect Treatment Comparison software, Ottawa, Canada). As expected, the nine pairwise ORs reported by Schulman in his Table 3 (1) are identical or nearly identical with one another, and the very slight variations affecting some of these estimates are likely the result of rounding rather than reflecting a true difference. As third point, in examining the question of whether relative or absolute measures should be incorporated in a NETMA, Schulman addresses a very relevant issue that a few weeks ago has been re-proposed to the attention of the scientific community (6). In line with the general recommendation of preferring absolute measures as opposed to relative ones, we have recently proposed to incorporate absolute measures (e.g. the number needed to treat) also in NETMA (7, 8). However, we admit that using these measures is only a small advancement in the framework of a technique that in turn generates only small advancements.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.258 | 0.662 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.010 | 0.009 |
| Bibliometrics | 0.010 | 0.010 |
| Science and technology studies | 0.001 | 0.005 |
| Scholarly communication | 0.012 | 0.015 |
| Open science | 0.007 | 0.005 |
| Research integrity | 0.010 | 0.013 |
| Insufficient payload (model declined to judge) | 0.024 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".