MétaCan
Menu
Back to cohort
Record W2057892004 · doi:10.4141/cjps08400

Why is MIXED analysis underutilized

2008· article· en· W2057892004 on OpenAlexaffvenueabout
Rong‐Cai Yang

Bibliographic record

VenueCanadian Journal of Plant Science · 2008
Typearticle
Languageen
FieldAgricultural and Biological Sciences
TopicGenetics and Plant Breeding
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsEnvironmental science

Abstract

fetched live from OpenAlex

The Operations Manual (http://pubs.nrc-cnrc.gc.ca/aicjournals/instruct/operations-manual.pdf) for the Canadian Journal of Plant Science, the Canadian Journal of Soil Science and the Canadian Journal of Animal Science (Revised 2007, page 17) states: ‘‘The GLM procedure of SAS has been widely used for analysis of variance; however, it was designed to analyze data having fixed effects only. Models that have both fixed and random effect should be analyzed using the MIXED procedure of SAS. This is also important in analyzing datasets with repeated observations on the same experimental unit that have heterogeneous variances over time and/or unequal within subject timedependent correlations.’’ Despite these clearly stated guidelines regarding the basic requirements and expectations of the statistical analysis, many submissions to the CJPS have continued the use of GLM when clearly MIXED should be used. A quick inspection of the first two 2008 issues of CJPS reveals that, of the 33 papers with a description of experimental designs and statistical analyses, eight used MIXED, 21 explicitly or implicitly indicated the use of GLM and the remaining four used other software (e.g., SPSS and GenStat). Both fixed and random effects are present in most of these studies judging from their description of experiments, but GLM rather than MIXED has been used in the majority of cases. Such trend of underutilization of MIXED is probably true as well in the previous volumes of CJPS and other agricultural journals. The purpose of this letter is to discuss causes and consequences of such underutilization in crop and agronomic research. Before embarking on such discussion, it is important to briefly review the concept of fixed and random effects. The determination of whether an effect is fixed or random in crop and agronomic studies is not always easy and has been debated in the scientific literature. In crop and agronomic experiments, treatments or combinations of treatments are often chosen intentionally and thus should be fixed effects. Moreover, these experiments are usually carried out at multiple sites and over several years to infer about the treatment performance for future years over a wide region. Such broader inference assumes that site and year effects are random, with sites being a random sample of all possible sites in the region and years being a random set of future years. However, these assumptions are rarely fulfilled in practice since the locations are not always randomly selected, and years may not be representative of future years (Steel et al. 1997). Despite the practical difficulty, both sites and years are generally considered as random for the broader inference. There are several reasons why GLM remains commonly used in the scientific literature for experiments with both fixed and random effects. First, it is often argued that GLM and MIXED give the same results when the data sets are balanced. It is true that obtaining a balanced data set is relatively easier in crop and agronomic experiments than in forestry and animal experiments, where factors such as tree mortality or cost of animals may make it more difficult to achieve the data balance. It is also true that with a balanced data set, the estimated variances of random effects would be identical whether the estimation procedure is the residual (restricted) maximum likelihood (REML) (the default method of MIXED) or TYPE I to Type IV of GLM, provided that these variance estimates are not negative. In this case, GLM would indeed provide the same F-tests of fixed effects if the random effects are specified in the RANDOM statement and the TEST option is added. If the true random effects are small and/or sample sizes are small, the negative variance estimates may be obtained using GLM. However any variance by definition should not be less than zero and negative estimates have no meaning. When this happens, REML (the default in MIXED) sets the negative variance estimates to zero regardless of whether or not the data set is balanced! Such different modes of handling negative variance estimates by GLM vs. MIXED would lead to different F-tests of the same fixed effects. Interestingly, the GLM vs. MIXED difference in F-tests due to the presence of negative variance estimates creates a new issue of which F-test should be used. In other words, should we use an F-test based on negative but unbiased variance components or an F-test based on nonnegative but biased variance components? This remains to be an open question even among statisticians (Littell et al. 2002). Nevertheless,

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.315
Threshold uncertainty score0.996

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.002
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.052
GPT teacher head0.195
Teacher spread0.143 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2008
Admission routes3
Has abstractyes

Explore more

Same venueCanadian Journal of Plant ScienceSame topicGenetics and Plant BreedingFrench-language works237,207