Assessing SDGs: A New Methodology to Measure Sustainability
Bibliographic record
Abstract
The FEEM project APPS – Assessment, Projections and Policy of Sustainable Development Goals – focuses on the quantitative assessment of the seventeen Sustainable Development Goals (SDGs), adopted by the United Nations at the end of September 2015. The project consists of two phases. The first, retrospective, computes indicators for all SDGs in 139 countries and then derives a composite multi-dimensional index and a worldwide ranking of current sustainability. This allows informing on strengths and weaknesses of today socio-economic development, as well as environmental criticalities, all around the world. The second phase, prospective, aims at evaluating the future trends of sustainability in the world by 2030. The assessment of the SDGs is carried out by means of an extended version of the recursive-dynamic computable general equilibrium ICES macro-economic model that includes social and environmental indicators. The final goal is to highlight future challenges left unsolved in the next 15 years of socio-economic development and to analyze costs and benefits of specific policies to support the achievement of proposed targets. This paper presents the methodology and the results of the retrospective assessment. Five main steps are described: i) screening of indicators eligible to address the UN SDGs; ii) data collection from relevant sources; iii) organization in the three pillars of sustainability (economy, society, environment); iv) normalization to a common metrics; v) aggregation of the 25 indicators in composite indices by pillars as well as in the multi-dimensional index. The final ranking summarizes countries’ sustainability performance. As expected, Middle-North European countries are at top of the ranking (Sweden, Norway and Switzerland the first three), with the most industrialized European countries such as Germany and UK, however, penalized by insufficient environmental performance. Other highly developed countries are between 24th (Canada) and 52nd place (United States). The emerging nations are scattered in our sustainability ranking. Brazil (43rd) and Russia (45th) precede China (80th) and India (102nd), the latter two especially penalized because of their social complexity. The worst performances, in terms of overall sustainability, are in Sub-Saharan Africa (Comoros, the Central African Republic and Chad occupy the last places in the ranking).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.006 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".