Multivariable Modeling and Multivariate Analysis for the Behavioral Sciences
Bibliographic record
Abstract
Multivariable Modeling and Multivariate Analysis for the Behavioral Sciences by Brian Everitt is a second-level applied statistics book which is aimed at those who need to build simple models in behavioural sciences. It provides varied sets of real world data so that the reader can gain insights into how these models are relevant to solving real life problems. The book starts with a chapter on data, measurement and models. Whereas most books treat the first chapter as a warm-up to what is to come, Everitt reminds us here of some very important but often neglected principles such as the limitations of significance tests, the importance of highlighting the aspects of the data that are relevant to the substantive arguments, the significance of experiments and the relevance of power in choosing the sample size. The exercises reinforce the principles discussed with the help of real world problems. The second chapter is a preliminary look at the data by using graphic methods. Here the author illustrates the use of less widely used graphics such as dot plots, leaf plots and boxplots, probability plots and various scatter plots followed by a brief discussion of how graphs can be used to mislead the reader. Although these techniques are widely known, the graphic capabilities of R make them much more accessible. The graphic techniques are discussed not so much in the context of presenting the data but in the context of making sense of the data. The next three chapters describe locally weighted linear regression, simple linear regression and its equivalence to analysis of variance. These are followed by logistic regression, survival analysis and linear mixed models for longitudinal analysis. The four subsequent chapters deal with the structure of multivariate data by using interdependent models such as principal components analysis, factor analysis and cluster analysis. The final chapter presents methods to analyse multivariate data drawn from several different populations. In his exposition, Everitt separates the technical aspects from the practical aspects of models. Because technical aspects are presented in self-contained sections, non-technical self-study readers can follow the material without becoming bogged down in formulae. Widely available statistical packages such as SAS, SPSS and Systat make it possible for non-technical readers to implement the models that are described. For those who do not have access to such packages, Everitt provides code in R language, which of course is free. For those who would like access to actual data so that they can practice what they have learnt, several data sets are made available from a companion Web site. Solutions to selected problems appear at the end of the book, which include R code to implement many of the techniques that are described in the book. Clarity and conciseness have always been the hallmarks of Everitt’s writing. This book is no exception. Anyone looking for a clearly written text on the subject that is also practitioner oriented needs to look no further.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.086 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.005 | 0.003 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.001 | 0.005 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.003 | 0.008 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".