Information for Authors: Is the Advice Regarding the Reporting of Residuals in Regression Analysis Incomplete? Should Cook’s Distance Be Included?
Bibliographic record
Abstract
If regression analysis is used for statistical evaluation of the data, authors must supply … standard deviations of residuals (Sy|x, often called standard errors of estimates)… Residuals plots [e.g., Bland-Altman] are often useful. —Extract from “Information for Authors” (2006) The Clinical Chemistry “Information for Authors” recommends that, when regression analysis is used, SDs of residuals must be supplied. (They are not always provided.) As Cook and Weisberg note (1), this conceptual approach dates back to the early 1960s, but by the late 1970s, attention was increasingly directed to assessing the influence of individual observations on the results of regression analysis. The concept of influence (or leverage) can be illustrated by 2 simple examples. In Fig. 1A1 , the regression line is shown for 4 in-line cases. When case 5 is added, the new regression line is slightly leveraged toward it (Fig. 1C1 ), but note that the case 5 residual is large (Fig. 1E1 ) and the regression lines are nearly parallel. However, when case 5 (Fig. 1B1 ) is added, the new regression line is much more influenced by its presence (Fig. 1D1 ). This case forces the regression line close to it, and its residual is correspondingly small (Fig. 1F1 ). What are the differences between these 2 cases? When an outlier is close to the mean value of x (as in case 5 in Fig. 1A1 ), its influence is small (Fig. 1C1 ), whereas when the outlier is a long way from the mean value of x and out of line of the initial regression line (as in case 5 in Fig. 1B1 ), its influence is large (Fig. 1D1 ). It will be noted that the respective values of the residuals do not reflect the effect of influence, because an influential case may decrease the magnitude of the residual. Supplemental Fig. 1 displays the same data with Deming regression (see Fig. 1 in the Data Supplement that accompanies the online version of this Opinion at http://www.clinchem.org/content/vol52/issue10). The Deming residuals mirror those shown in Fig. 11 , although they are of greater magnitude. Again, no relationship exists between the residuals and leverage. It should be noted that several alternatives to Cook’s distance have been proposed (3)(6)(7), although for various reasons they are less favored. Stuart et al. (3) recommend that one or other of these indices should always be examined. The generation of Cook’s distance values for a data set might seem daunting, but it should be realized that this capability is available in many statistical programs, such as SPSS, SAS, Minitab, S-Plus, R (8), and Arc(9)(10). The latter 2 programs are freely available. It is of interest that the 3 statistical programs with clinical chemistry applications (Analyze-it, MedCalc, and CBstat) do not (yet) provide this capability. The Deming Cook’s distance equivalent is obtained by replacing ri by rdemi (Eq. 8), although software for calculating these values is not currently available. I have chosen to use a medical data set (11), used by Altman in his textbook (12), to illustrate the factors that affect the values of Cook’s distance (see Figs. 2 and 3 in the online Data Supplement). Reporting only residual values (which the journal requires) does not always identify cases that influence the resulting regression line. The more thorough approach, with the use of Cook’s distance (or its equivalents), provides much more insight regarding the regression model. Many current texts illustrate the use and value of estimating Cook’s distance in linear regression (9)(13)(14)(15). Accordingly, I propose that the provision of Cook’s distance values or other similar measure should be encouraged when regression analysis is reported. Such a suggestion is particularly valuable when the sample size is small. Dr. Henderson died on June 29, 2006. The effect of an outlier on the resulting least squares regression line. Solid lines show the regression line with all 5 cases. (A), when case 5 outlier is near to the mean of the x values. (B), when case 5 outlier is distant from the mean of the x values. (C), leverage values for all 5 cases in A. (D), leverage values for all 5 cases in B. (E), crude residual values for all 5 cases in A. (F), crude residual values for all 5 cases in B. I thank Drs. Sanford Weisberg and Dennis Cook (both at the University of Minnesota) and the technical support staff at Insightful Corp. for helpful advice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.041 | 0.491 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.005 | 0.009 |
| Open science | 0.005 | 0.003 |
| Research integrity | 0.008 | 0.006 |
| Insufficient payload (model declined to judge) | 0.240 | 0.308 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".