A Comprehensive Study on Quality Aspects and Industry Perspective in Backporting
Bibliographic record
Abstract
Code quality assurance has emerged as a well-established pillar to ensure effective software development and maintenance. Large and intricate software systems (e.g., Linux, Adobe Reader, Collabora, etc.) often require developers to version source code and manage multiple stable releases simultaneously. In general, prior stable releases are maintained by Backporting, which refers to applying changes taken from a newer release to an old release. In such cases, software quality assurance can be challenging for old stable releases for a variety of reasons. First, older releases often differ from newer (upstream) releases in terms of code dependency, infrastructure, architecture, and maintenance strategy. Second, backport maintenance is often not reviewed consistently to prioritize upstream development. Lastly, as software releases evolve with con- tinuous maintenance, developers often fail to maintain code in accordance with essential quality standards. As a result, poor-quality code and design choices can emerge in stable releases, affecting their reliability and stability. Although researchers have extensively scrutinized the quality assurance practice and paradigm in upstream versions, how code evolves and quality issues arise in the backporting procedure is yet to be ex- plored. Thus, to fill this gap in existing research work, we analyzed how software quality deviates in terms of size, complexity, and coupling throughout backporting. We found that as software releases evolve, backports are responsible for 11.5% and 12.3% of quality degradation and improvement, respectively. Furthermore, backporting can significantly affect the complexity and size of old releases. In our second study, we strive to observe when and why technical debt (fine-grained quality issues) arises as stable releases evolve through backporting. Our exploration reveals that the early-phase of Apache releases and the mid-phase of Eclipse and Python releases are more prone to technical debts in the release life cycle. Moreover, we found develop- ers’ high workload and low exposure can lead to new technical debts in stable releases. Lastly, in our third study, we explored expert opinions to get a non-biased, reliable view of developers’ needs and challenges to ensure code quality in the backporting process. We asked 38 experts and analyzed their challenges to ensure code quality for the backporting process. This study reveals the several subjective factors behind the failed quality assessment in backporting practice, including code comprehension, lack of efficient decision-making standards, lack of testing guidelines, scarcity of organization and tool support, etc.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.034 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.005 | 0.008 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.005 | 0.010 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".