Amer: A New Attribute-Missing Network Embedding Approach
Bibliographic record
Abstract
Network embedding which aims to learn a low dimensional representation of nodes is a powerful technique for network analysis. While network embedding for networks with complete attributes has been widely investigated, in many real-world applications the attributes of partial nodes are unobserved (i.e., missing) due to privacy concern or resource limit. Very recently, several network embedding methods have been proposed for attribute-missing networks. They first complete the missing attributes and then use the complemented network to learn network embedding. The parameters of these two processes cannot be adjusted by each other, resulting in compromised results. To address this problem, we propose a unified model in which the process of completing missing attributes and the process of learning embedding are not separated but closely intertwined. Being specific, completing missing attributes is under the guidance of learning network representation via mutual information maximization, and the complemented attributes directly enter network representation module which will generate further feedback for completing missing attributes. We further impose attribute-structure relationship constraint for completing missing attributes by designing a new generative adversarial networks (GANs) model. To the best of our knowledge, this is the first unified model for attribute-missing network embedding. Empirical results on real-world datasets show the superiority of our new method over other state-of-the-art methods on four network analysis tasks, including node classification, node clustering, link prediction, and network visualization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".