Housing safety and health academic and public opinion mining from 1945 to 2021: PRISMA, cluster analysis, and natural language processing approaches
Bibliographic record
Abstract
Housing safety and health problems threaten owners' and occupiers' safety and health. Nevertheless, there is no systematic review on this topic to the best of our knowledge. This study compared the academic and public opinions on housing safety and health and reviewed 982 research articles and 3,173 author works on housing safety and health published in the Web of Science Core Collection. PRISMA was used to filter the data, and natural language processing (NLP) was used to analyze emotions of the abstracts. Only 16 housing safety and health articles existed worldwide before 1998 but increased afterward. U.S. scholars published most research articles (30.76%). All top 10 most productive countries were developed countries, except China, which ranked second (16.01%). Only 25.9% of institutions have inter-institutional cooperation, and collaborators from the same institution produce most work. This study found that most abstracts were positive (n = 521), but abstracts with negative emotions attracted more citations. Despite many industries moving toward AI, housing safety and health research are exceptions as per articles published and Tweets. On the other hand, this study reviewed 8,257 Tweets to compare the focus of the public to academia. There were substantially more housing/residential safety (n = 8198) Tweets than housing health Tweets (n = 59), which is the opposite of academic research. Most Tweets about housing/residential safety were from the United Kingdom or Canada, while housing health hazards were from India. The main concern about housing safety per Twitter includes finance, people, and threats to housing safety. By contrast, people mainly concerned about costs of housing health issues, COVID, and air quality. In addition, most housing safety Tweets were neutral but positive dominated residential safety and health Tweets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".