A comparison between visual estimates and image analysis measurements to determine septoria leaf blotch severity in winter wheat
Bibliographic record
Abstract
Methods to estimate disease severity vary in accuracy, reliability, ease of use and cost. Severity of septoria leaf blotch ( SLB , caused by Z ymoseptoria tritici ) was estimated by four raters and by image analysis (assumed actual values) on individual leaves of winter wheat in order to explore accuracy and reliability of estimates, and to ascertain whether there were any general characteristics of error. Specifically, the study determined: (i) the accuracy and reliability of visual assessments of SLB over the full range of severity from 0 to 100%; (ii) whether certain 10% ranges in actual disease severity between 0 and 100% were more prone to estimation error compared with others; and (iii) whether leaf position affected accuracy within those ranges. Lin's concordance correlation analysis of all severities (0–100%) demonstrated that all raters had estimates close to the actual values (agreement: ρ c = 0·92–0·99). However, agreement between actual SLB severities and estimates by raters was less good when compared over short 10% subdivisions within the 0–100% range (ρ c = −0·12 to 0·99). Despite common rater imprecision at estimating low and high SLB severities, individual raters differed considerably in their accuracy over the short 10% subdivisions. There was no effect of leaf position on accuracy or precision of severity estimate on separate leaves (L1–L3). Pursuing efforts in understanding error in disease estimation should aid in improving the accuracy of assessments, making visual estimates of disease severity more useful for research and applied purposes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".