Understanding Recent Trends in Global Sustainable Development Goal 6 Research: Scientometric, Text Mining and an Improved Framework for Future Research
Bibliographic record
Abstract
The fulfilment of Sustainable Development Goal (SDG) 6, concerning water and sanitation, is critical in itself and also conditional for the other 16 SDGs being met. The purpose of this study was to understand the scientific research trajectories, spatiotemporal development, scientific collaboration, ongoing research themes, and gaps related to SDG 6. We propose a coupling of bibliometric and text mining methods in this work, to statistically portray the impact of water research on the accomplishment of SDG 6. Through the Web of Science database, we focused on a single UN SDG goal (i.e., six related publications that were current (2015–2021)). The study was performed on the chosen 289 publications. With the analysis of Keywords Plus, abstracts, titles, as well as author keywords, we looked at the performance of authors, publications, journals, institutions, and nations in terms of publishing. To obtain an insight into the water and sanitation study topic, we used co-citation, co-occurrence, cooperation networks, theme networks and cluster analysis, word dynamics, thematic evolution, and other techniques. We filtered out five distinguishing themes using text mining and showed their temporal trends. The main outcome is that participation, as well as collaboration with countries of the Global South, is still lacking in the SDG 6 research sphere. Therefore, as an insight from this study, we proposed a conceptual framework, the sustainable development of water and sanitation (SDWS) framework, to classify the research domain of water and sanitation regarding its connections to the environment, economy, and society (i.e., sustainable development). The scientometric and text analysis results provide the contemporary state and overview of the water and sanitation research field, whereas the second, conceptual framework section, provides a better understanding of qualitative contents, by revealing the insights gained, as well as the important work to be done in future water and sanitation studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.012 |
| Science and technology studies | 0.004 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".