Mechanisms of self-organized criticality in social processes of knowledge creation
Bibliographic record
Abstract
In online social dynamics, a robust scale invariance appears as a key feature of collaborative efforts that lead to new social value. The underlying empirical data thus offers a unique opportunity to study the origin of self-organized criticality (SOC) in social systems. In contrast to physical systems in the laboratory, various human attributes of the actors play an essential role in the process along with the contents (cognitive, emotional) of the communicated artifacts. As a prototypical example, we consider the social endeavor of knowledge creation via Questions and Answers (Q&A). Using a large empirical data set from one of such Q&A sites and theoretical modeling, we reveal fundamental characteristics of SOC by investigating the temporal correlations at all scales and the role of cognitive contents to the avalanches of the knowledge-creation process. Our analysis shows that the universal social dynamics with power-law inhomogeneities of the actions and delay times provides the primary mechanism for self-tuning towards the critical state; it leads to the long-range correlations and the event clustering in response to the external driving by the arrival of new users. In addition, the involved cognitive contents (systematically annotated in the data and observed in the model) exert important constraints that identify unique classes of the knowledge-creation avalanches. Specifically, besides determining a fine structure of the developing knowledge networks, they affect the values of scaling exponents and the geometry of large avalanches and shape the multifractal spectrum. Furthermore, we find that the level of the activity of the communities that share the knowledge correlates with the fluctuations of the innovation rate, implying that the increase of innovation may serve as the active principle of self-organization. To identify relevant parameters and unravel the role of the network evolution underlying the process in the social system under consideration, we compare the social avalanches to the avalanche sequences occurring in the field-driven physical model of disordered solids, where the factors contributing to the collective dynamics are better understood.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".