Near-miss software clones in open source games: An empirical study
Bibliographic record
Abstract
Developers tend to reuse source code by copy/paste. This form of reuse introduces code clones to software systems. Cloning in games can happen in different levels of granularity. The extreme case is known as Game Clone where the complete project is being cloned, e.g., by making a new independent branch which the original game's source code constitutes the seed for the new branch. Although it has been more than two decades since research on code clones started, various characteristics of cloning in open source games has not been studied. Therefore, there is no specific evidence on status of cloning e.g., dominant clone type in games. In this paper, we present an empirical study on code cloning in open source games by applying a state of the art clone detector, NiCad, and a clone visualization and analysis tool, VisCad. We identify both exact and near-miss clones from more than twenty open source C, Java and Python-based games, such as Jake2, Duo, Hexen2, jChess and OpenRPG from five different categories. Furthermore, we analyze a set of metrics for code clones in several different dimensions, including language, category, clone density and clone location to answer six essential research questions about the current status of cloning in open source games. Our research illustrates that cloning happens not only at inter-project level but also intra-project, specifically in First Person Shooter games developed using C (Type-1 50% clones). Such observation concretely shows the necessity of adopting clone management systems for game development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".