Notice bibliographique
Résumé
Thank you for the opportunity to respond to the letter sent by Dr Dickey (Dickey, 2003) in which he included comments on my paper on the pitfalls in the design and analysis of efficacy trials in subfertility (Daya, 2003). I have several points to make. First, it appears that Dr Dickey is confusing the analytical and inferential phases of clinical trials. It is well known that the purpose of a clinical trial on efficacy evaluation is to estimate the true (but unknown) effect of the experimental intervention when compared with the control intervention. In this context, the larger the sample size the more precisely will the researchers be able to obtain this estimate of truth. The analytical part of the trial is directed towards determining how likely (or not) the estimate observed in the trial is the result of a chance finding; the lower the likelihood (probability), the lower the play of chance and the higher the likelihood that what was observed represents a true finding. Next, the magnitude of the effect of the intervention is addressed. Owing to sampling variability, the size of the treatment effect will vary considerably. Hence, larger trials, performed repeatedly, are more likely to provide a robust estimate of the size of the treatment effect. These issues of statistical significance and size of treatment effect are addressed by the analytical step in the clinical trial. In the inferential step, the researcher tries to make sense of the results as they apply to the population from which the sample was derived, i.e. what do the results of the study indicate about the effect of treatment in the population? What conclusions can one make about the value of the experimental intervention in the patients, identified by the inclusion criteria. Thus, statistical significance (an analytical aspect) has to be differentiated from clinical significance or relevance (an inferential aspect) of efficacy evaluation by clinical trials. The former relies on strict methodological criteria, whereas the latter requires individual judgment. These two aspects are not mutually exclusive, as Dr Dickey implies, but rather complementary. The second point of apparent misunderstanding by Dr Dickey is the concept of ‘unit of analysis’. In his letter, Dr Dickey uses the term ‘unit of analysis’ to mean outcome measure (e.g. live birth, pregnancy and so on), rather than the unit (or subject) that was randomized, which is the accepted definition. The experimental intervention is administered to the subject (selected randomly so that bias can be minimized) and compared with the control intervention administered to another subject (also selected randomly). The outcomes are then compared between the two groups of subjects (i.e. the unit of analysis is the subject). Thus, in ART trials of two interventions, it is the subject (i.e. female patient) who is being randomized and so the outcome must pertain to her (as is the case with pregnancy, live birth and even number of oocytes retrieved). The use of implantation rate as an outcome measure is inappropriate because the denominator is no longer the total number of subjects (the unit of analysis) but the total number of embryos transferred. However, because the embryos were not randomly allocated they cannot be used as the unit of analysis. The third point, also related to the unit of analysis, pertains to the sample size that is calculated on the basis of assumptions made about the expected event rates (such as pregnancy rates) in the experimental and control groups. Whenever the sample studied is much larger than required, there is a higher likelihood that smaller differences in outcome events between the two groups will become statistically significant. Thus, switching the focus away from pregnancy rate to implantation rate, the denominator chosen is no longer the subject but the number of embryos transferred, a number that is much larger than the number of subjects studied. It is not surprising, therefore, that one finds significant differences in implantation rates (embryos as denominator) even though pregnancy rates (subject as denominator) are not significantly different. This fact, no doubt, is the reason why many investigators include this outcome when comparing treatments in ART cycles. The fourth point pertains to the outcome measures that should be used. Depending on the objective of the trial, as identified by the research question, the investigator will choose a primary outcome (on which the sample size calculation is based) and one or more secondary outcomes. In the field of infertility treatment, because there is no consensus on the key outcome measures that should be utilized, several outcomes have been reported, including ovulation rate, pregnancy rate (both clinical and ongoing), live birth rate, numbers of oocytes and so on. The focus for most patients, investigators and other stakeholders is pregnancy, and until consensus can be reached on how this outcome should be reported, it may be prudent to provide data on all three relevant pregnancy outcomes viz clinical pregnancy, ongoing pregnancy and live birth rate (Daya, 2003). Such reporting also implies that data will be provided on numbers of miscarriage, ectopic pregnancy and stillbirth. The concern about the high rates of multiple pregnancy can be addressed, in part, by selecting singleton live birth as the outcome of choice (Daya, 2003). This initiative is supported by the increasing call for single embryo transfer in women undergoing treatment with ART and should become the standard for efficacy evaluation studies. In summary, useful and valid inferences can be made only when trials are executed with high methodological quality and rigour, and analysed with the necessary attention to detail, as it pertains to the randomized subject who is the unit of analysis, using appropriate and clinically relevant outcome measures.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,124 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».