High-Throughput Phenotyping and Genomic Prediction in Multi-Environment Plant Breeding Field Trials
Bibliographic record
Abstract
Plant breeding requires phenotyping for multiple traits throughout the growing season in large, multilocation field trials. This large-scale phenotyping produces the data necessary to select lines with a desirable combination of traits. The use of unoccupied aerial vehicles (UAVs) equipped with sensors to assist breeding programs in field data collection offers potential to overcome challenges with manual phenotyping. UAV-based imaging also offers opportunities to monitor and measure field trials in ways not previously possible. However, approaches to use image data most effectively to support field trial evaluation are still lacking. In this thesis, a diverse hexaploid wheat (Triticum aestivum L.) nested-association mapping population consisting of 1160 recombinant inbred lines was evaluated in yield trials conducted at three locations during the 2020 and 2021 growing seasons. UAV-based multispectral imaging was conducted at 10-15 timepoints throughout phenological development and spectral summary statistics, spectral indices, and texture features were extracted at each timepoint. LASSO regression models trained on image feature sets were able to predict days to heading (mean R2 = 0.76), days to maturity (mean R2 = 0.84), plant height (mean R2 = 0.70), and grain yield (mean R2 = 0.64) within testing environments more accurately than a gradient boosted decision tree, simple linear regression, and spatial models. Cross-environment prediction was conducted, and higher grain yield prediction accuracies were observed for LASSO regression models using image features (mean R2 = 0.34) than genomic prediction models alone (mean R2 = 0.26). However, the best cross-environment grain yield predictions were observed by combining image feature and genomic prediction models (mean R2 = 0.39). Image-based prediction modeling was also applied to durum wheat (Triticum turgidum L. var durum) breeding population field trials of over 2600 yield plots evaluated in four environments. Within-environment LASSO regression prediction accuracies of up to R2 = 0.88 were observed, indicating the potential for high-throughput phenotyping of complex traits in breeding populations. Genome-wide association mapping of image features was performed and significant marker-trait associations for all features were identified. The texture feature Energy was highlighted as detecting a marker-trait association near the locus of the wheat height gene Rht-B1. Visual inspection of plot images revealed Energy detected the presence of lodging. This thesis provides insight into the potential application of high-throughput phenotyping and genomic prediction to improve the evaluation of wheat breeding field trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".