Temporal learning for dynamic graph
Bibliographic record
Abstract
A graph is a data structure to model a complex system of entities connected in a particular relation.These entities are nodes in the graph, and the connections are edges.A dynamic graph is a graph that evolves in its nodes, edges or both.We can use it to analyze networks that are not static, such as social networks, academic citation networks and city traffic networks.The dynamic graph is a widely used data structure in various domains.However, the exploration of machine learning with dynamic graphs is still in its early stage.What are the drivers of a dynamic graph's evolution?How can we learn the temporal information from a dynamic graph's history?How can we determine if we should use dynamic graph learning algorithms to analyze a given graph?These are still open questions that are well worth to explore.In this thesis, I try to address the research questions above by a survey of recently developed supervised dynamic graph learning algorithms and proposing a dynamic graph temporal learning framework.Based on the framework above, I conducted an initial study on measuring the significance of temporal patterns to prepare for the research in predicting performance gain from dynamic graph learning algorithms.Dr.Xue Liu, I learned from him how to do research from a high-level point of view, how to discover and focus on the high-impact research ideas, how to plan our work more efficiently, and many other best practices in conducting scientific research.All of these are very beneficial to my future career.The collaboration system he established between his students and other research groups also opened up my eyesight to different scientific domains.I am genuinely thankful to my supervisor Dr.Xue Liu.If there is one thing I regret during my study at McGill, that would be my giving up on the research idea that Dr.Kieran O'Donnell helped me establish.I selected to work on a start-up idea instead of continuing that potential high-impact work.I must say sorry to Dr.Kieran O'Donnell for not continuing to work with him because of my greediness.I also need to say 'thank you' to him for all the favours he did for me.This regretful experience taught me that I should not be greedy for money,
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.021 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.007 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".