MétaCan
Menu
Retour à la cohorte
Enregistrement W4249082046 · doi:10.1093/humrep/deh021

Reply

2003· article· en· W4249082046 sur OpenAlexaff
S. Daya

Notice bibliographique

RevueHuman Reproduction · 2003
Typearticle
Langueen
DomaineMathematics
ThématiqueStatistical Methods in Clinical Trials
Établissements canadiensMcMaster University
Organismes subventionnairesnon disponible
Mots-clésMedicineGynecology

Résumé

récupéré en direct d'OpenAlex

Thank you for the opportunity to respond to the letter sent by Dr Dickey (Dickey, 2003) in which he included comments on my paper on the pitfalls in the design and analysis of efficacy trials in subfertility (Daya, 2003). I have several points to make. First, it appears that Dr Dickey is confusing the analytical and inferential phases of clinical trials. It is well known that the purpose of a clinical trial on efficacy evaluation is to estimate the true (but unknown) effect of the experimental intervention when compared with the control intervention. In this context, the larger the sample size the more precisely will the researchers be able to obtain this estimate of truth. The analytical part of the trial is directed towards determining how likely (or not) the estimate observed in the trial is the result of a chance finding; the lower the likelihood (probability), the lower the play of chance and the higher the likelihood that what was observed represents a true finding. Next, the magnitude of the effect of the intervention is addressed. Owing to sampling variability, the size of the treatment effect will vary considerably. Hence, larger trials, performed repeatedly, are more likely to provide a robust estimate of the size of the treatment effect. These issues of statistical significance and size of treatment effect are addressed by the analytical step in the clinical trial. In the inferential step, the researcher tries to make sense of the results as they apply to the population from which the sample was derived, i.e. what do the results of the study indicate about the effect of treatment in the population? What conclusions can one make about the value of the experimental intervention in the patients, identified by the inclusion criteria. Thus, statistical significance (an analytical aspect) has to be differentiated from clinical significance or relevance (an inferential aspect) of efficacy evaluation by clinical trials. The former relies on strict methodological criteria, whereas the latter requires individual judgment. These two aspects are not mutually exclusive, as Dr Dickey implies, but rather complementary. The second point of apparent misunderstanding by Dr Dickey is the concept of ‘unit of analysis’. In his letter, Dr Dickey uses the term ‘unit of analysis’ to mean outcome measure (e.g. live birth, pregnancy and so on), rather than the unit (or subject) that was randomized, which is the accepted definition. The experimental intervention is administered to the subject (selected randomly so that bias can be minimized) and compared with the control intervention administered to another subject (also selected randomly). The outcomes are then compared between the two groups of subjects (i.e. the unit of analysis is the subject). Thus, in ART trials of two interventions, it is the subject (i.e. female patient) who is being randomized and so the outcome must pertain to her (as is the case with pregnancy, live birth and even number of oocytes retrieved). The use of implantation rate as an outcome measure is inappropriate because the denominator is no longer the total number of subjects (the unit of analysis) but the total number of embryos transferred. However, because the embryos were not randomly allocated they cannot be used as the unit of analysis. The third point, also related to the unit of analysis, pertains to the sample size that is calculated on the basis of assumptions made about the expected event rates (such as pregnancy rates) in the experimental and control groups. Whenever the sample studied is much larger than required, there is a higher likelihood that smaller differences in outcome events between the two groups will become statistically significant. Thus, switching the focus away from pregnancy rate to implantation rate, the denominator chosen is no longer the subject but the number of embryos transferred, a number that is much larger than the number of subjects studied. It is not surprising, therefore, that one finds significant differences in implantation rates (embryos as denominator) even though pregnancy rates (subject as denominator) are not significantly different. This fact, no doubt, is the reason why many investigators include this outcome when comparing treatments in ART cycles. The fourth point pertains to the outcome measures that should be used. Depending on the objective of the trial, as identified by the research question, the investigator will choose a primary outcome (on which the sample size calculation is based) and one or more secondary outcomes. In the field of infertility treatment, because there is no consensus on the key outcome measures that should be utilized, several outcomes have been reported, including ovulation rate, pregnancy rate (both clinical and ongoing), live birth rate, numbers of oocytes and so on. The focus for most patients, investigators and other stakeholders is pregnancy, and until consensus can be reached on how this outcome should be reported, it may be prudent to provide data on all three relevant pregnancy outcomes viz clinical pregnancy, ongoing pregnancy and live birth rate (Daya, 2003). Such reporting also implies that data will be provided on numbers of miscarriage, ectopic pregnancy and stillbirth. The concern about the high rates of multiple pregnancy can be addressed, in part, by selecting singleton live birth as the outcome of choice (Daya, 2003). This initiative is supported by the increasing call for single embryo transfer in women undergoing treatment with ART and should become the standard for efficacy evaluation studies. In summary, useful and valid inferences can be made only when trials are executed with high methodological quality and rigour, and analysed with the necessary attention to detail, as it pertains to the randomized subject who is the unit of analysis, using appropriate and clinically relevant outcome measures.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,004
score de la tête « metaresearch » (Gemma)0,124
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Méthodes · Signal consensuel: aucune
Score de désaccord entre enseignants0,653
Score d'incertitude au seuil0,951

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0040,124
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0010,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,690
Tête enseignante GPT0,586
Écart entre enseignants0,104 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2003
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueHuman ReproductionMême sujetStatistical Methods in Clinical TrialsTravaux en français237 207