Network Analysis of The Evolution of an Open Source Development Community
Bibliographic record
Abstract
This research investigated the evolution of an open source development community by undertaking network analysis across time.Several open source communities have been previously studied using network analysis techniques, including Apache, Debian, Drupal, Python, and SourceForge.However, only static snapshots of the network were considered by previous research.In this research, we created a tool that can help researchers and practitioners to find out how the Eclipse development community dynamically evolved over time.The input dataset was collected from the Eclipse Foundation, and then the evolution of the Eclipse development community was visualized and analyzed by using different network analysis techniques.Six network analysis techniques were applied: (i) visualization, (ii) weight filtering, (iii) degree centrality, (iv) eigenvector centrality, (v) betweenness centrality, and (vi) closeness centrality.Results include the benefits of performing multiple techniques in combination, and the analysis of the evolution of an open source development community.Open source software: The distribution terms of open-source software must comply with the following criteria: (i) free redistribution (ii) program must include source code, (iii) derived works among committers, (iv) integrity of the author's source code, (v) no discrimination against persons or groups, (vi) no discrimination against fields of endeavour, (vii) license must not be specific to a product, (viii) license must not restrict other software, (ix) license must be technology-neutral.(Open Source Initiative, 2013) Project: A Project is the main operational unit at Eclipse.Specifically, all open source software development at Eclipse occurs within the context of a Project.Eclipse Projects are organized hierarchically.A special type of Project, Top-Level Projects, sits at the top of the hierarchy.Each Top-Level Project contains one or more Projects.A Top-Level Project is said to be the "parent" of those Projects.A Project that has a parent is oftentimes referred to as a Sub-Project (Eclipse
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.008 | 0.006 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.003 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".