Sustainability assessment in agriculture: emerging issues in voluntary sustainability standards and their governance
Bibliographic record
Abstract
Over the past two decades, voluntary sustainability standards (VSS) have emerged as instruments to improve social and environmental practices and to communicate sustainability standards in trade and business. However, debates about the correct assessment methodology for VSS risk causing duplication, overlaps, and fragmentation, undermining the value of VSS in sustainability transition. In this paper I propose materiality, theory of change and reflexive governance as the three building blocks of an appropriate framework for VSS and other sustainability assessment schemes in the food and agricultural sectors. Materiality is the specific criteria for defining and assessing factors that matter for sustainability, such as indicators, metrics, and rankings. Materiality is a process of social construction that enables stakeholder engagement and integrated knowledge production, going beyond just benchmarking entities against their competitors using standardized measures. Theory of change is a method that explains how interventions lead to desired outcomes and changes but is much more than a linear logic model of inputs and outputs. It sheds light on underlying assumptions, embedded contexts and long- and short-term dynamics. Reflexive evaluation consists of a single-loop process that follows a problem-detection-correction course and double- and triple-loop learning that allows assumptions and learned propositions to be challenged. It highlights unintended outcomes and offers alternatives to conducting interventions, which is different from conventional monitoring and evaluation methods, which focus on measuring the attainment of intended outcomes only. The study concludes that the semantic meaning of “standards” and “assessment” in agricultural VSS abstract the complex nature of sustainability because of overly linear meaning. In the complex construct of sustainability assessment, the role of VSS is not to conclude a success or a failure but to encourage knowledge-based learning and accountable governance because social change is an open-ended validation and adaption process. The framework proposed by this paper offers a solution by calling for integrated knowledge production resulting from interdisciplinary assessments and learning-oriented actions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.076 | 0.065 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.008 |
| Science and technology studies | 0.005 | 0.084 |
| Scholarly communication | 0.025 | 0.033 |
| Open science | 0.004 | 0.012 |
| Research integrity | 0.010 | 0.014 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".