Avoiding Mistakes with Component Selection in Requirements Engineering
Bibliographic record
Abstract
Modern software systems are increasingly being built by combining preexisting building blocks, which are called components, rather than developing everything from scratch [267]. This component-based software development (CBSD) approach can accelerate delivery and reduce cost, but its success is highly dependent on the ability of developers to select appropriate components [133]. Poor component selection can lead to integration failures, reduced main-tainability, security vulnerabilities, and higher costs [157, 355]. Existing research has proposed several methods, tools, and metrics, yet these solutions have struggled to gain traction in industry [157]. Therefore, despite the importance of component selection, current practices remain largely informal and ad hoc, with limited tool support [23, 27]. Developers continue to face uncertainty when evaluating the quality, compatibility, and long-term sustainability of components. At the same time, advances in artificial intelligence (AI) present new opportunities to help developers navigate vast sources of information about components such as GitHub (GH) issues, forums, and documentation [23]. This thesis investi-gates how AI, and in particular natural language processing (NLP), can support developers in assessing components to make more informed selection decisions. The research was carried out in three phases. First, the problem was defined through a literature survey that synthesized recent work on component selection methods, tools, and evaluation criteria, highlighting the strengths, weaknesses, and gaps that limit adoption in practice. Second, the problem was further explored through a survey with almost 100 devel-opers and CBSD experts, which revealed expectations and challenges of practitioners, as well as a strong demand for automated support and evidence-based evaluation of components. Third, a potential solution was introduced in the form of a prototype approach, which applied two subsets of NLP, natural language inference (NLI) and large language models (LLMs), to analyze discussions around components and automatically present pros and cons related to a set of selected quality criteria. The effectiveness of this approach was then evaluated through a second survey involving 37 developers, who evaluated its usefulness and limitations. The results showed that while AI techniques have promise in providing developers with richer and more contextual information, their current limitations in accuracy, transparency, and trustworthiness must be carefully managed. The research contributes an updated review of the literature on component selection practices, insight into the needs of the practitioner from a survey study, and an initial validation of AI-driven techniques for component evaluation. The findings highlight the potential for future work to integrate AI-driven selection support into practical developer workflows, with an eye on applicability, explainability, and performance. This work makes three concrete contributions to the field of component-based software engineering (CBSE). First, it provides a comprehensive literature review that organizes exist-ing selection methods and catalogs over 700 quality attributes into a structured taxonomy, identifying gaps such as the lack of automated component selection tools and overlooked factors like licensing and community health. Second, it offers empirical insights from a sur-vey of nearly 100 practitioners, revealing which criteria developers prioritize (e.g., reliability, documentation, security) and highlighting common pitfalls in current practice, as well as a strong demand for better tool support. Third, it introduces and evaluates a novel AI-driven approach to component evaluation. A curated dataset of 271 developer discussions was labeled for key quality attributes, and a prototype tool was built using NLI and LLM techniques to automatically extract and present evidence (pros and cons) from public sources. An initial user study with 37 developers demonstrated the prototype’s potential to reduce evaluation effort, while also underscoring the need for transparency and explainability in AI-assisted tools. Together, these contributions advance the understanding of component selection challenges and lay the groundwork for integrating intelligent support into developer workflows, ultimately aiming to make CBSD more informed, reliable, and efficient.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.073 | 0.186 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.003 | 0.005 |
| Scholarly communication | 0.006 | 0.011 |
| Open science | 0.003 | 0.007 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".