MétaCan
Menu
Retour à la cohorte
Enregistrement W2129454670 · doi:10.1373/clinchem.2006.068296

Information for Authors: Is the Advice Regarding the Reporting of Residuals in Regression Analysis Incomplete? Should Cook’s Distance Be Included?

2006· article· en· W2129454670 sur OpenAlexaff
A R Henderson

Notice bibliographique

RevueClinical Chemistry · 2006
Typearticle
Langueen
DomaineMathematics
ThématiqueAdvanced Statistical Methods and Models
Établissements canadiensWestern University
Organismes subventionnairesnon disponible
Mots-clésOutlierRegression analysisRegressionResidualStatisticsLeverage (statistics)Linear regressionStandard deviationSimple linear regressionRegression diagnosticRegression toward the meanMathematicsEconometricsStandard errorLine (geometry)Polynomial regressionAlgorithm

Résumé

récupéré en direct d'OpenAlex

If regression analysis is used for statistical evaluation of the data, authors must supply … standard deviations of residuals (Sy|x, often called standard errors of estimates)… Residuals plots [e.g., Bland-Altman] are often useful. —Extract from “Information for Authors” (2006) The Clinical Chemistry “Information for Authors” recommends that, when regression analysis is used, SDs of residuals must be supplied. (They are not always provided.) As Cook and Weisberg note (1), this conceptual approach dates back to the early 1960s, but by the late 1970s, attention was increasingly directed to assessing the influence of individual observations on the results of regression analysis. The concept of influence (or leverage) can be illustrated by 2 simple examples. In Fig. 1A1 , the regression line is shown for 4 in-line cases. When case 5 is added, the new regression line is slightly leveraged toward it (Fig. 1C1 ), but note that the case 5 residual is large (Fig. 1E1 ) and the regression lines are nearly parallel. However, when case 5 (Fig. 1B1 ) is added, the new regression line is much more influenced by its presence (Fig. 1D1 ). This case forces the regression line close to it, and its residual is correspondingly small (Fig. 1F1 ). What are the differences between these 2 cases? When an outlier is close to the mean value of x (as in case 5 in Fig. 1A1 ), its influence is small (Fig. 1C1 ), whereas when the outlier is a long way from the mean value of x and out of line of the initial regression line (as in case 5 in Fig. 1B1 ), its influence is large (Fig. 1D1 ). It will be noted that the respective values of the residuals do not reflect the effect of influence, because an influential case may decrease the magnitude of the residual. Supplemental Fig. 1 displays the same data with Deming regression (see Fig. 1 in the Data Supplement that accompanies the online version of this Opinion at http://www.clinchem.org/content/vol52/issue10). The Deming residuals mirror those shown in Fig. 11 , although they are of greater magnitude. Again, no relationship exists between the residuals and leverage. It should be noted that several alternatives to Cook’s distance have been proposed (3)(6)(7), although for various reasons they are less favored. Stuart et al. (3) recommend that one or other of these indices should always be examined. The generation of Cook’s distance values for a data set might seem daunting, but it should be realized that this capability is available in many statistical programs, such as SPSS, SAS, Minitab, S-Plus, R (8), and Arc(9)(10). The latter 2 programs are freely available. It is of interest that the 3 statistical programs with clinical chemistry applications (Analyze-it, MedCalc, and CBstat) do not (yet) provide this capability. The Deming Cook’s distance equivalent is obtained by replacing ri by rdemi (Eq. 8), although software for calculating these values is not currently available. I have chosen to use a medical data set (11), used by Altman in his textbook (12), to illustrate the factors that affect the values of Cook’s distance (see Figs. 2 and 3 in the online Data Supplement). Reporting only residual values (which the journal requires) does not always identify cases that influence the resulting regression line. The more thorough approach, with the use of Cook’s distance (or its equivalents), provides much more insight regarding the regression model. Many current texts illustrate the use and value of estimating Cook’s distance in linear regression (9)(13)(14)(15). Accordingly, I propose that the provision of Cook’s distance values or other similar measure should be encouraged when regression analysis is reported. Such a suggestion is particularly valuable when the sample size is small. Dr. Henderson died on June 29, 2006. The effect of an outlier on the resulting least squares regression line. Solid lines show the regression line with all 5 cases. (A), when case 5 outlier is near to the mean of the x values. (B), when case 5 outlier is distant from the mean of the x values. (C), leverage values for all 5 cases in A. (D), leverage values for all 5 cases in B. (E), crude residual values for all 5 cases in A. (F), crude residual values for all 5 cases in B. I thank Drs. Sanford Weisberg and Dennis Cook (both at the University of Minnesota) and the technical support staff at Insightful Corp. for helpful advice.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,041
score de la tête « metaresearch » (Gemma)0,491
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: Présentation des résultats · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: aucune
Score de désaccord entre enseignants0,959
Score d'incertitude au seuil0,804

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0410,491
Méta-épidémiologie (sens strict)0,0010,002
Méta-épidémiologie (sens large)0,0030,002
Bibliométrie0,0060,007
Études des sciences et des technologies0,0020,002
Communication savante0,0050,009
Science ouverte0,0050,003
Intégrité de la recherche0,0080,006
Charge utile insuffisante (le modèle a refusé de juger)0,2400,308

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,263
Tête enseignante GPT0,523
Écart entre enseignants0,260 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeThéorique ou conceptuel
DomainePrésentation des résultats
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations11
Publié2006
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueClinical ChemistryMême sujetAdvanced Statistical Methods and ModelsTravaux en français237 207