Nullius in Verba1: Advancing Data Transparency in Industrial Ecology
Bibliographic record
Abstract
Summary With the growth of the field of industrial ecology (IE), research and results have increased significantly leading to a desire for better utilization of the accumulated data in more sophisticated analyses. This implies the need for greater transparency, accessibility, and reusability of IE data, paralleling the considerable momentum throughout the sciences. The Data Transparency Task Force (DTTF) was convened by the governing council of the International Society for Industrial Ecology in late 2016 to propose best‐practice guidelines and incentives for sharing data. In this article, the members of the DTTF present an overview of developments toward transparent and accessible data within the IE community and more broadly. We argue that increased transparency, accessibility, and reusability of IE data will enhance IE research by enabling more detailed and reproducible research, and also facilitate meta‐analyses. These benefits will make the results of IE work more timely. They will enable independent verification of results, thus increasing their credibility and quality. They will also make the uptake of IE research results easier within IE and in other fields as well as by decision makers and sustainability practitioners, thus increasing the overall relevance and impact of the field. Here, we present two initial actions intended to advance these goals: (1) a minimum publication requirement for IE research to be adopted by the Journal of Industrial Ecology ; and (2) a system of optional data openness badges rewarding journal articles that contain transparent and accessible data. These actions will help the IE community to move toward data transparency and accessibility. We close with a discussion of potential future initiatives that could build on the minimum requirements and the data openness badge system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.293 | 0.562 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.010 | 0.009 |
| Science and technology studies | 0.006 | 0.012 |
| Scholarly communication | 0.028 | 0.028 |
| Open science | 0.005 | 0.027 |
| Research integrity | 0.009 | 0.018 |
| Insufficient payload (model declined to judge) | 0.023 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".