Avoiding Mistakes with Component Selection in Requirements Engineering
Notice bibliographique
Résumé
Modern software systems are increasingly being built by combining preexisting building blocks, which are called components, rather than developing everything from scratch [267]. This component-based software development (CBSD) approach can accelerate delivery and reduce cost, but its success is highly dependent on the ability of developers to select appropriate components [133]. Poor component selection can lead to integration failures, reduced main-tainability, security vulnerabilities, and higher costs [157, 355]. Existing research has proposed several methods, tools, and metrics, yet these solutions have struggled to gain traction in industry [157]. Therefore, despite the importance of component selection, current practices remain largely informal and ad hoc, with limited tool support [23, 27]. Developers continue to face uncertainty when evaluating the quality, compatibility, and long-term sustainability of components. At the same time, advances in artificial intelligence (AI) present new opportunities to help developers navigate vast sources of information about components such as GitHub (GH) issues, forums, and documentation [23]. This thesis investi-gates how AI, and in particular natural language processing (NLP), can support developers in assessing components to make more informed selection decisions. The research was carried out in three phases. First, the problem was defined through a literature survey that synthesized recent work on component selection methods, tools, and evaluation criteria, highlighting the strengths, weaknesses, and gaps that limit adoption in practice. Second, the problem was further explored through a survey with almost 100 devel-opers and CBSD experts, which revealed expectations and challenges of practitioners, as well as a strong demand for automated support and evidence-based evaluation of components. Third, a potential solution was introduced in the form of a prototype approach, which applied two subsets of NLP, natural language inference (NLI) and large language models (LLMs), to analyze discussions around components and automatically present pros and cons related to a set of selected quality criteria. The effectiveness of this approach was then evaluated through a second survey involving 37 developers, who evaluated its usefulness and limitations. The results showed that while AI techniques have promise in providing developers with richer and more contextual information, their current limitations in accuracy, transparency, and trustworthiness must be carefully managed. The research contributes an updated review of the literature on component selection practices, insight into the needs of the practitioner from a survey study, and an initial validation of AI-driven techniques for component evaluation. The findings highlight the potential for future work to integrate AI-driven selection support into practical developer workflows, with an eye on applicability, explainability, and performance. This work makes three concrete contributions to the field of component-based software engineering (CBSE). First, it provides a comprehensive literature review that organizes exist-ing selection methods and catalogs over 700 quality attributes into a structured taxonomy, identifying gaps such as the lack of automated component selection tools and overlooked factors like licensing and community health. Second, it offers empirical insights from a sur-vey of nearly 100 practitioners, revealing which criteria developers prioritize (e.g., reliability, documentation, security) and highlighting common pitfalls in current practice, as well as a strong demand for better tool support. Third, it introduces and evaluates a novel AI-driven approach to component evaluation. A curated dataset of 271 developer discussions was labeled for key quality attributes, and a prototype tool was built using NLI and LLM techniques to automatically extract and present evidence (pros and cons) from public sources. An initial user study with 37 developers demonstrated the prototype’s potential to reduce evaluation effort, while also underscoring the need for transparency and explainability in AI-assisted tools. Together, these contributions advance the understanding of component selection challenges and lay the groundwork for integrating intelligent support into developer workflows, ultimately aiming to make CBSD more informed, reliable, and efficient.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,073 | 0,186 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,006 | 0,005 |
| Études des sciences et des technologies | 0,003 | 0,005 |
| Communication savante | 0,006 | 0,011 |
| Science ouverte | 0,003 | 0,007 |
| Intégrité de la recherche | 0,002 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».