Making Sense of the Minefield of Footprint Indicators
Bibliographic record
Abstract
In recent years, footprint indicators have emerged as a popular mode of reporting environmental performance. The prospect\nis that these simplified metrics will guide investors, businesses, public sector policymakers and even consumers of everyday\ngoods and services in making decisions which lead to better environmental outcomes. However, without a common “DNA”,\nthe ever expanding lexicon of footprints lacks coherence and may even report contradictory results for the same subject matter.1\nThe danger is that this will ultimately lead to policy confusion and general mistrust of all environmental disclosures. Footprints are especially interesting metrics because they seek to express the environmental performance of products and organizations from a life cycle perspective. The life cycle perspective is important to avoid misleading claims based only on a selected life cycle stage. For example, the water used to manufacture beverages may be important, but if a beverage includes sugar, irrigation water used to cultivate sugar cane could be a greater concern. The focus on environmental performance distinguishes footprints from technical efficiency measures, such\nas energy use efficiency or water use efficiency, which typically only make sense when applied to a single life cycle stage as they lack local environmental context. However, unlike technical efficiency, which can usually be accurately measured and verified, footprint indicators, with their wider view of environmental performance, are usually calculated using models which can differ in scope, complexity and model parameter settings. Despite the noble intention of using footprints to evaluate and report environmental performance, the potential inconsistency between different approaches acts as a deterrent to use in many public policymaking and business contexts and can lead to confusing and contradictory messages in the marketplace.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.008 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".