The impact of natural constraints in linear regression of log transformed response variables
Bibliographic record
Abstract
Abstract In linear regression, log transforming the response variable is the usual workaround regarding departures from the assumption of normality. However, the response variable is often subject to natural constraints, which can result in a truncated distribution of the residual errors on the log scale. In forestry, allometric relationships and tree growth are two typical examples a natural constraint; the response variable cannot be negative. Traditional least squares estimators do not account for constrained response variables. For this study, a modified maximum likelihood (MML) estimator that takes natural constraints into account was developed. This estimator was tested through a simulation study and showcased with black spruce tree diameter increment data. Results show that the ordinary least squares estimator underestimated large conditional expectations of the response variable on the original scale. In contrast, the MML estimator showed no evidence of bias for large sample sizes. Departures from distributional assumptions cannot be overlooked when the model is used for predictive purposes. Both Monte Carlo error propagation and prediction intervals rely on these assumptions. In this context, the MML estimator developed for this study can be used to properly propagate the errors and produce reliable prediction intervals.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.042 | 0.190 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".