Assessing the efficacy and toxicity of leflunomide: Every little counts!
Bibliographic record
Abstract
Dear Sir, I read with interest the original article, the editorial based on it and the letter to the Editor which subsequently followed about leflunomide published in the Journal recently. Several important issues need further examination. Firstly, in the quest to separate toxicity from the efficacy no single approach is likely to be adequate. The letter suggests “meta-analysis is a step forward”. Like RCTs the meta-analysis are not without biases which can originate from the methodology. Moreover there could be a difficulty in comparing studies from the different eras. To overcome the later difficulty we in our own two widely cited meta-analyses of therapy for psoriatic arthritis and of the toxicity of glucocorticoids in RA have used patient withdrawal as a marker for efficacy (less likely to be withdrawn if therapy is effective) and toxicity (more likely to be withdrawn if therapy is causing harm). It’s important to note that leflunomide turned out to be only second to anti-TNF inhibitors in benefit versus risk with the ratio of numbers needed to treat (NNT) to numbers needed to harm (NNH) of 0.45 and 0.25 respectively. Never the less even if the Quality of Reporting of Meta-analyses (QUOROM) guidelines are followed verbatim the meta-analyses are only as good as the studies it analyses. I therefore believe that something like the wide ranging set of EULAR initiatives for the glucocorticoids are more appropriate in the assessment of the toxicity of leflunomide. In the Indian context, before longitudinal observational databases are put forward as the panacea for the shortcomings of the current void in the published literature regarding the harm and benefit in real life situation, it would be prudent to establish that all such databases are universally recording high quality reliable data which can be collated to pave the way for its use as the evidence which may change the practice. Probably well intentioned audits are not at all
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.039 | 0.159 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.006 | 0.012 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.010 | 0.027 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".