Semi-Automated Roadside Image Data Collection for Characterization of Agricultural Land Management Practices
Bibliographic record
Abstract
Land cover management practices, including the adoption of cover crops or retaining crop residue during the non-growing season, has important impacts on soil health. To broadly survey these practices, a number of remotely sensed products are available but issues with cloud cover and access to agriculture fields for validation purposes may limit the collection of data over large regions. In this study, we describe the development of a mobile roadside survey procedure for obtaining ground reference data for the remote sensing of agricultural land use practices. The key objective was to produce a dataset of geo-referenced roadside digital images that can be used in comparison to in-field photos to measure agricultural land use and land cover associated with crop residue and cover cropping in the non-growing season. We found a very high level of correspondence (>90% level of agreement) between the mobile roadside survey to in-field ground verification data. Classification correspondence was carried out with a portion of the county-level census image data against 114 in-field manually categorized sites with a level of agreement of 93%. The few discrepancies were in the differentiation of residue levels between 30–60% and >60%, both of which may be considered as achieving conservation practice standards. The described mobile roadside image capture system has advantages of relatively low cost and insensitivity to cloudy days, which often limits optical remote sensing acquisitions during the study period of interest. We anticipate that this approach can be used to reduce associated field costs for ground surveys while expanding coverage areas and that it may be of interest to industry, academic, and government organizations for more routine surveys of agricultural soil cover during periods of seasonal cloud cover.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".