Self-Liking Group in Networks With Multi-Class Nodes
Bibliographic record
Abstract
Nodes in complex networks are generally allocated into groups using community detection methods. These communities are based on the interactions between nodes (links). Conversely, in machine learning, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">clustering methods</i> group data points into classes based on their attribute's similarities regardless of their interactions. Although both communities and clustering methods classify data points into groups, they are fundamentally different. Clustering relies on attribute similarity, while communities focus on interaction patterns. The present study bridges these two distinct approaches by introducing a new concept - Self-Liking Groups (SLG). Based on entropy considerations, SLG quantifies the preference of node classes to interact with similar ones based on their communication patterns, thus combining both the community and the clustering methods. We demonstrate SLG in three case studies: (i) A career network of 2.5 million companies, linked by 8 million job switches. Here, SLG reveals the openness of different industrial sectors to workers in other sectors. For example, the Healthcare sector shows the highest SLG, i.e., it is the least open to accepting workers from other sectors, while the Energy sector has a high SLG, but only for educated workers. Also, managers' shift between different sectors is more limited due to higher SLG. (ii) A scientific co-authorship network where SLG measures the openness of collaboration between different countries. China, India and Japan, have stronger SLG and are thus more likely to collaborate with scientists in their own country compared to the USA, Canada, and most EU countries. (iii) In the medical scientific research space, SLG reveals that Japan, a country known for its longevity, is extremely close compared to China or India. We also find that SLG is a stable measure across various community detection methods and initial parameter spaces. This implies that SLG captures a fundamental property of networks with heterogeneous nodes and is useful in analyzing real complex network scenarios.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".