MétaCan
Menu
Back to cohort
Record W3015042751 · doi:10.1037/met0000265

Partitioning variation in multilevel models for count data.

2020· article· en· W3015042751 on OpenAlexaff
George Leckie, William J. Browne, Harvey Goldstein, Juan Merlo, Peter C. Austin

Bibliographic record

VenuePsychological Methods · 2020
Typearticle
Languageen
FieldMathematics
TopicStatistical Methods and Bayesian Inference
Canadian institutionsUniversity of Toronto
FundersEconomic and Social Research Council
KeywordsOverdispersionCategorical variableIntraclass correlationStatisticsCount dataMultilevel modelMathematicsCluster analysisNegative binomial distributionPoisson distributionItem response theoryVariance (accounting)Binary dataEconometricsCorrelationBinary numberPsychometrics

Abstract

fetched live from OpenAlex

A first step when fitting multilevel models to continuous responses is to explore the degree of clustering in the data. Researchers fit variance-component models and then report the proportion of variation in the response that is due to systematic differences between clusters. Equally they report the response correlation between units within a cluster. These statistics are popularly referred to as variance partition coefficients (VPCs) and intraclass correlation coefficients (ICCs). When fitting multilevel models to categorical (binary, ordinal, or nominal) and count responses, these statistics prove more challenging to calculate. For categorical response models, researchers appeal to their latent response formulations and report VPCs/ICCs in terms of latent continuous responses envisaged to underly the observed categorical responses. For standard count response models, however, there are no corresponding latent response formulations. More generally, there is a paucity of guidance on how to partition the variation. As a result, applied researchers are likely to avoid or inadequately report and discuss the substantive importance of clustering and cluster effects in their studies. A recent article drew attention to a little-known exact algebraic expression for the VPC/ICC for the special case of the two-level random-intercept Poisson model. In this article, we make a substantial new contribution. First, we derive exact VPC/ICC expressions for more flexible negative binomial models that allows for overdispersion, a phenomenon which often occurs in practice. Then we derive exact VPC/ICC expressions for three-level and random-coefficient extensions to these models. We illustrate our work with an application to student absenteeism. (PsycInfo Database Record (c) 2020 APA, all rights reserved).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.040
metaresearch head score (Gemma)0.143
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.040
Threshold uncertainty score0.212

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0400.143
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0030.006
Bibliometrics0.0040.007
Science and technology studies0.0020.004
Scholarly communication0.0060.007
Open science0.0070.007
Research integrity0.0030.008
Insufficient payload (model declined to judge)0.0120.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.644
GPT teacher head0.585
Teacher spread0.059 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations67
Published2020
Admission routes1
Has abstractyes

Explore more

Same venuePsychological MethodsSame topicStatistical Methods and Bayesian InferenceFrench-language works237,207