MétaCan
Menu
Back to cohort
Record W7135003010

Exploring Cross-Language Software Similarity Analysis Using Source Code Context

2025· article· en· W7135003010 on OpenAlexfundno aff
Kawser Wazed Nafi

Bibliographic record

VenueUniversity Library (University of Saskatchewan) · 2025
Typearticle
Languageen
FieldComputer Science
TopicSoftware Engineering Research
Canadian institutionsnot available
FundersNatural Sciences and Engineering Research Council of CanadaCanada First Research Excellence Fund
KeywordsSoftware developmentSource codeSoftware constructionStatic program analysisCode reviewSoftware maintenanceDocumentationSoftware qualityCode reuseSoftware portability
DOInot available

Abstract

fetched live from OpenAlex

The rapid growth of multi-language and cross-platform software development has created an urgent need for effective techniques to identify functional similarity across programming languages. Developers routinely reuse or reimplement functionally similar code blocks across cross-language and multilingual software systems, resulting in intentional and unintentional cross-language similar code fragments, as well as the adaptation of APIs and libraries that serve similar purposes but are implemented in different languages. While these adaptations can improve portability and broaden software reach to various users, they also increase development cost, maintenance complexity, and the potential for inconsistency or defects. Despite recent advances in machine learning, code representation learning, and Large Language Models (LLMs), existing approaches for cross-language software similarity often struggle with deep syntactic reasoning, diverse coding styles, and limited availability of high-quality multi-lingual code datasets. This thesis is grounded in the premise that accurate detection of cross-language code similarity can significantly mitigate longstanding challenges in cross-language software development and maintenance. Motivated by this premise, the thesis investigates the foundational problem of establishing reliable, robust cross-language code-similarity measures. The proposed investigation aims to support a wide range of software engineering tasks, including single-language, cross-language, and multi-language development and maintenance activities. Drawing on a comprehensive systematic literature review, the thesis identifies key limitations in the state of the art and proposes five complementary contributions across four levels of code granularity. First, it introduces a universal software similarity detector (CroLSim) that categorizes cross-language software applications by leveraging API call documentation similarity. Second, it presents a source code feature-driven and API documentation-adapted cross-language clone detection model (CLCDSA) that combines syntactic features with API documentation semantics similarity to identify cross-language clones more accurately. Third, it develops an LLM-guided, multimodal framework (XLCoCo) that fuses multi-intent source code information retrieval from LLMs and attention-based VAEs to predict structural feature similarity, improving the performance of cross-language code-to-code search and clone detection tasks. Fourth, it proposes XLibRec, a technique for recommending analogical cross-language libraries by mining reliable library usage information from different developer community discussion forums, along with Library short descriptions collected from various package managers. Finally, it introduces XAPIRec, an efficient method for analogical API mapping based on API usage patterns, mined from functionally equivalent API usage patterns collected in an automatic way, and LLM-driven API document similarity, which completely replaces the need to manually mine functionally similar parallel code fragments or any prior knowledge of true mapped API or a labeled API mapping dataset. Together, these contributions form a scalable ecosystem that advances automation, accuracy, and practical applicability in industry-level cross-language software development and maintenance. The techniques are extensively evaluated against state-of-the-art baselines across diverse datasets and programming languages, demonstrating consistent improvements in precision, recall, ranking quality, and real-world usability. Overall, this thesis offers a unified framework to support developers and organizations in building, understanding, and maintaining robust cross-language software systems in large scale.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.060
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.023
Threshold uncertainty score0.037

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0070.060
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0230.013
Science and technology studies0.0010.001
Scholarly communication0.0040.008
Open science0.0020.004
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.028
GPT teacher head0.223
Teacher spread0.195 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueUniversity Library (University of Saskatchewan)Same topicSoftware Engineering ResearchFrench-language works237,207