Omnia munda mundis (‘to the pure, all things are pure’)
Notice bibliographique
Résumé
We are grateful for the opportunity to reply and comment on the Taggart et al. [1] and Freemantle et al. [2] pieces. We believe that a detailed description of the background and context is necessary for better understanding by the readers. Our individual patient meta-analysis of the relationship between different coronary artery bypass grafting (CABG) conduits and outcomes published in the European Journal of Cardio-Thoracic Surgery (EJCTS) [3] is based on a pooled dataset of individual patient data from 4 contemporary CABG trials that was initially established by all key investigators of each study, as well as CABG experts and statistical and methodological editors of major cardiac surgery journals, to evaluate sex differences in outcomes after CABG surgery [4]. Cornell Medicine in New York was established as the analytic unit based on the previous successful experiences with CABG individual patient data meta-analyses. The planned work on sex differences in CABG outcomes was successfully completed and published in the European Heart Journal [5]. Using the same methodological process, a comparison of the outcomes of different CABG conduits was selected by the group as the next project. When the results of this second analysis were shared with the investigator group, Dr Taggart (who was an author of the European Heart Journal manuscript) decided to withdraw his name as an author of this second project. The manuscript on CABG conduits was accepted for publication in the EJCTS, where the Statistical Reviewer noted: ‘The design of the study is sophisticated and coherent with the aim. […] The results are robust and clearly reported’. Dr Taggart was invited by EJCTS to write an Editorial Comment where he raised issues regarding the published work and called for an independent re-analysis of the data. In the name of transparency and open science, we agreed with the EJCTS request to share our dataset and codes for the purpose of a double-check (omnia munda mundis). Taggart et al. [1] raise the following points in their commentary: The mean age of the left internal thoracic artery (LITA) and radial artery (RA) group was younger than the other groups: this is incorrect. In the unmatched population, the bilateral internal thoracic artery (BITA) group was significantly younger {63.6 years [interquartile range (IQR) 57.2–69.8]} compared to the other 2 groups [LITA + RA: 65 years (IQR 59.6–70.3); LITA + saphenous vein graft (SVG): 66 years (IQR 60–72); see Supplementary Material, Table S2 in our manuscript [3]]. After matching, there were no substantial differences in population age among the 3 groups, as described by a standardized mean difference (SMD)<0.1. Patients in the ‘LITA ± RA’ group received a lower number of grafts than in the other groups: this is also incorrect. The number of grafts for the patients included in the matched cohorts was similar across groups and the median number of grafts was 3 (IQR 3–4) in all 3 groups. The statistical model should account for the heterogeneity of treatment effects across trials: we agree on this, and this is what we did. In fact, our propensity score model included the trial identifier as a covariable and, therefore, the trial was considered a variable to be balanced across the matched groups. Similarly, our sensitivity analyses were based on mixed-effect Cox models which included the trial identifier as a random intercept. It is interesting to note that this approach was not used by Freemantle et al. [2] in their independent analysis, highlighting differential analytic approaches to the same research questions. Matching with replacement requires adjustment of standard errors: we agree on this, and this is what we did. In our analysis, standard errors were in fact adjusted to account for matching with replacement. Standard errors were calculated as clustered robust standard errors, which account for both the paired nature of the observations and the matching weights. Handling of missing covariates: our analysis was a complete case analysis, with no missing data. Apparent exclusion of the Radial Artery Patency and Clinical Outcomes (RAPCO) trial: RAPCO patients were excluded from the propensity score-matched analysis due to the high rate of missing baseline data which would have not allowed the application of a non-parsimonious model to calculate the propensity score. This is standard practice in propensity score-based comparisons. However, the findings of RAPCO are consistent with our analyses, showing better outcomes for the RA compared to the BITA and the GSV [6], and therefore, it is reasonable to expect a strengthening of our results if RAPCO had been included. It must also be noted that RAPCO was included in our sensitivity analysis using multivariable Cox regression, which was consistent with our main analysis and was robust to even strong confounding effects based on the E-values. In summary, some of the issues raised by Taggart et al. [1] were factually incorrect and others are easily addressed by a clarification of the methodology used. The request for an independent re-analysis of a published work is generally justified by concerns of data fabrication (not appliable here as all the included trials were previously published) or evidence of analytic errors (such as individual subcategories not adding up to the total, unplausible variance of baseline characteristics, etc.); Taggart’s editorial did not provide any of those and, thus, the request for a re-analysis of our dataset was not justified. We are not aware of any other work that was the subject to an independent re-analysis after publication based on similar concerns. However, we decided to comply with the EJCTS request because of our strong commitment to open science and, most importantly, to establish a clear precedent of transparency and trust in the cardiac surgery literature (omnia munda mundis). We are appreciative that the findings from our original analysis could be reproduced by Freemantle et al. [2] and that our published results are not in question. We also notice that Freemantle et al. proceeded to perform additional analyses that were not part of our manuscript. Of note, the statistical analysis plan and codes used for this EJCTS independent analysis were not shared with us or with any of the individual trial authors, nor were details of the peer-review process for the Freemantle et al. manuscript. While the results of Freemantle et al. additional analysis are quantitively different than ours, they are qualitatively consistent with our model. The observed quantitative differences can be easily explained on the basis of a key difference in the analytic approach. Propensity score matching is based on the principle of finding patients in the experimental group that closely match corresponding patients in the control group; the matching algorithm is driven by the characteristics of patients in the control group and sequentially evaluates patients in the experimental group to identify the best possible match based on a number of pre-defined variables. The selection of the control group is critical and has important implications in this case: the EJCTS independent analysis and our analysis used a different control group and, by doing so, included different patient populations and answered different research questions. Our analysis uses the SVG as the control group to answer the question of what would have happened to patients who received the SVG if they had received the RA or BITA instead. As SVG is the most commonly used conduit for CABG, this is the key question in CABG surgery—the SVG is extremely easy to use and most CABG surgeons use SVG on the majority of their patients, so the SVG-matched population is representative of the majority of CABG patients operated by the majority of CABG surgeons. Conversely, the EJCTS independent analysis uses the BITA as the control group and answers the question of what would have happened to patients who received BITA if they had received RA or SVG. The use of the BITA is technically more complex than the use of RA or SVG, especially in patients with diffuse and distal coronary disease. In clinical practice, BITA is routinely used only by few dedicated surgeons in a minority of highly selected cases; in fact, the proportion of patients receiving BITA is small, both in our EJCTS work in question and in other large contemporary database [7]. BITA patients generally also have better conduits and more favourable coronary anatomy compared to patients who received other strategies, although those data are not captured in most clinical databases and cannot be effectively adjusted using statistical techniques. In other words, the BITA-matched population included by Freemantle et al. is made of a small and highly selected group of healthier CABG patients operated by highly experienced CABG surgeons. In this patient population, the treatment effect of any conduit intervention is likely to be small, due to the unique (more favourable) baseline profile and also to the higher technical skills of the operators. The difference in the included patient population is the most likely reason for the differences in results between our analyses and the EJCTS independent analysis. While for the BITA versus SVG comparison, findings from all the models consistently show no difference between conduits and for the BITA versus RA comparison, the differences are quantitative, not qualitative (the point estimate is in favour of RA even in the EJCTS additional analysis); differences in baseline patient characteristics (mostly unmeasured) and sample size (our comparison is based on 1776 pairs, while the EJCTS independent analysis is based on 595 pairs) are the most likely explanations for the observed quantitative variation in RA effect. To address Freemantle et al. sophisticated concerns of possible imbalances even in the presence of reassuring SMDs, we have replicated the EJCTS pairwise (as opposed to triplets) propensity score-matched approach, but we have used the SVG (rather than BITA) as the control group, as this makes the most clinical sense. By doing so have replicated the level of balance achieved by Freemantle et al. (Table 1) and we have found results that are consistent with our triplets analysis (RA better than BITA and SVG, no difference between BITA and SVG—see Fig. 1). Results of 3 analytic approaches estimating the association between conduits used and mortality. BITA: bilateral internal thoracic artery; CI: confidence interval; HR: hazard ratio; RA: radial artery; SV: saphenous vein. Baseline characteristics in the pairwise matched groups BITA: bilateral internal thoracic artery; CVA: cerebrovascular accidents; IQR: interquartile range; LVEF: left ventricle ejection fraction; MI: myocardial infarction; NYHA: New York Heart Association; PTCA: percutaneous transluminal coronary angioplasty; PVD: peripheral vascular disease; RA: radial artery; SMD: standardized mean difference; SV: saphenous vein. Baseline characteristics in the pairwise matched groups BITA: bilateral internal thoracic artery; CVA: cerebrovascular accidents; IQR: interquartile range; LVEF: left ventricle ejection fraction; MI: myocardial infarction; NYHA: New York Heart Association; PTCA: percutaneous transluminal coronary angioplasty; PVD: peripheral vascular disease; RA: radial artery; SMD: standardized mean difference; SV: saphenous vein. There are several important key messages from this exercise in open science. From a clinical perspective, the main research question is if in patients undergoing CABG the use of the RA or the BITA as compared to SVG improves clinical outcomes. As CABG is the most common cardiac surgery procedure (>400 000/year in the USA only) and SVG is used in ∼90% of the cases, the question is relevant. The reading of the totality of the analyses is that there is strong evidence that BITA is no better than SVG. This is also consistent with available randomized data from the large Arterial Revascularization Trial [8] as well as contemporary patency studies [9]. It seems possible that the suboptimal BITA results are not related to the biology of the conduits, but rather to the surgical technique and the deliverability of the operation; however, there are no data to support this theory. For the RA versus SVG comparison, the EJCTS independent analysis relied on the transitivity property and did not perform a direct comparison, so there is less information; we note that, as the RA is the second most commonly used complementary CABG conduit after the SVG [7], a direct comparison of the RA and SVG would have been important and in our opinion more important than the presented RA versus BITA comparison (a secondary question, as only a minority of CABG patients can receive the RA and BITA interchangeably). The available randomized data support better patency and clinical outcomes for RA versus SVG [6], providing strong biologic and clinical support to our findings. We agree with Taggart et al. that the RA treatment effect in our analysis is likely exaggerated by unadjusted confounders, and this is clearly acknowledged in the long limitation section of our manuscript. However, the RA effect is qualitatively consistent with the 30% mortality reduction with RA versus BITA at 15 years recently reported by the RAPCO trial and, despite all the described limitations, cannot be ignored or dismissed as implausible. It is widely accepted that different analytic models applied to the same dataset may provide quantitative different results [10]; while qualitative differences are much less frequent, statistical significance based on the classical 0.05 threshold may vary between models and this may generate confusion. Pre-specification, and ideally registration, of the analysis plan is important even in observational studies to avoid cherry-picking of the models; in addition, sensitivity analyses should be performed to confirm the solidity of observational findings. In our EJCTS piece in question, the analysis plan was pre-specified and Cox regression and E-values were used to confirm the propensity-matching findings. While we applaud the EJCTS efforts in assuring the rigour of published data, more work is needed to better address similar issues in the future. It is important that the rules and terms for this type of data-sharing requests are mutually agreed upfront to avoid potentially dangerous conflicts. Pre-definition of the re-analysis plan in conjunction with individual trial teams (possibly different than the authors of the original manuscript) and clear rules for publication, authorship, and peer-review of the re-analysis are important points that should be pre-defined and openly shared by all journals, ideally with the involvement and support of the International Committee of Medical Journal Editors and the World Association of Medical Editors. Most importantly, data must be read and interpreted with open mind, and results that go against personal beliefs and biases must not be dismissed as wrong or biologically implausible. This is the foundation of science and research progress (omnia munda mundis). The issue of the effect of using different conduits for CABG has been debated for 5 decades and the answer to date is still based on few studies with important limitations. Our ongoing Randomized comparison of the clinical Outcome of single vs Multiple Arterial grafting (ROMA trial—NCT03217006) will provide a solid answer in the next few years.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,021 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,005 | 0,003 |
| Communication savante | 0,004 | 0,005 |
| Science ouverte | 0,001 | 0,002 |
| Intégrité de la recherche | 0,032 | 0,039 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,010 | 0,008 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».