A multidimensional statistical model for wood data analysis, with density estimated from CT scanning data as an example
Bibliographic record
Abstract
The trunk of a tree can be seen as a spatiotemporal sampling domain from the statistical perspective, where space is represented by direction horizontally and height vertically, and time through annual growth rings. In this framework, wood properties such as density can be the object of data collection for given estimation and testing purposes. We present a multidimensional statistical model, the tensor normal distribution, in which the variation (variance) of and dependency (covariance) between wood property measurements made for different years at various locations in a tree trunk can be inferred. Its application requires a smaller number of replicates (trees) than the traditional vector normal distribution because variances and covariances for directions and growth rings, for example, must be the same at all heights, up to a multiplicative constant. This assumption on the variance–covariance structure is called “separability”, and we explain how to test it. An illustration with wood density estimates obtained from computed tomography scanning data for 11 white spruce ( Picea glauca (Moench) Voss) trees is presented. This example is completed by assessing differences in mean wood density according to location in the trunk, with analysis-of-variance F-tests adjusted for the estimated variances and covariances obtained by fitting the model.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.035 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".