A New Web-Based Big Data Analytics for Dynamic Public Opinion Mapping in Digital Networks on Contested Biotechnology Fields
Bibliographic record
Abstract
The expression "public opinion" has long been part of common parlance. However, its value as a scientific measure has been the topic of abundant academic debates over the past several decades. Such debates have produced more variety and contestations rather than consensus on the very definition of public opinion, let alone on how to measure it. This study reports on the usefulness of web-based big data digital network analytics in deciphering the distributed meanings and sense making related to controversial biotechnology applications. Using stem cell therapies as a case study, we argue that such digital network analysis can complement the traditional opinion polls while avoiding the sampling bias that is typical of opinion polls. Although the polls cannot account for the opinion dynamics, combining them with web-based big data analysis can shed light on three dimensions of public opinion essential for sense making: counts or volume of opinion data, content, and movement of opinions. This approach is particularly promising in the case of ongoing scientific controversies that increasingly overflow into the public sphere morphing into public political debates. In particular, our study focuses as a case study on public controversies over the clinical provision of stem cell therapies. Using web entities specifically addressing stem cell issues, including their dynamic aggregation, the internal architecture of the web corpus we report in this study brings the third dimension of public opinion (movement) into sharper focus. Notably, the corpus of stem cell networks through web connectivity presents hot spots of distributed meaning. Large-scale surveys conducted on these issues, such as the Eurobarometer of Biotechnology, reveal that European citizens only accept research on stem cells if they are highly regulated, while the stem cell digital network analysis presented in this study suggests that distributed meaning is promise centeredness. Although major scientific journals and companies tend to structure public opinion networks, our finding of promise centeredness as a key ingredient of distributed meaning and sense making is consistent with therapeutic tourism that remains as an important facet of the stem cell community despite the lack of material standards. This new approach to digital network analysis has crosscutting corollaries for rethinking the notion of public opinion, be it in electoral preferences or as we discuss in this study, for new ways to measure, monitor, and democratically govern emerging technologies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".