Development of a scaling factors framework to improve the approximation of software functional size with cosmic - iso19761
Bibliographic record
Abstract
Many software development organizations strive to deliver high-quality products while keeping a balance between customer satisfaction, time, and budget. The estimation of the effort of software development projects is one of the major challenges of these software development organizations. This challenge is typically faced at early phases of the software development life cycle. To tackle this challenge, the software development organizations use early estimation techniques to obtain early effort estimates (i.e. a priori estimates) in order to help project managers and technical leaders in projects planning and management. One of the methodologies for a priori effort estimation is based on the approximation of the expected software functionality. This requires the use of a measurement method to quantify this functionality: the literature refers to the measurement of the functional size of software products—including business applications. Various international standards have been adopted to measure the functional size of software such as ISO 19761: COSMIC. However, during the early phases of the software development life cycle, and more specifically in the approximation of the functional size of the software expected to be developed, the lack of detailed and complete software requirements specifications is common, which leads to many challenges. For instance, the level of granularity (i.e. the level of details) of the functional requirements specifications of software is identified subjectively using intuition, experience and/or opinions of the field experts. Also, there is no standardized notation to define a standard set of scaling factors to be assigned by the requirements engineers to the functional requirements specifications of software projects to identify theirs levels of granularity. These challenges affect the quality of the functional size approximation of software development projects, since the result of the functional size approximation process is one of the primary inputs for the a priori effort estimation process. These challenges prevent the estimators of software projects from building realistic effort estimation models. The motivation of this research project is to help software organizations and in particular projects managers and technical leaders to build more accurate effort estimation models by improving one of the inputs for the effort estimation process, in order to improve the planning, the management, and the development of software at early phases of the software development life cycle. The goal of this research project is to improve one of the inputs of the a priori effort estimation process, and in particular the functional size approximation of software development projects. The main research objective is to design a framework—to be used by the requirements engineers—that assigns scaling factors to early versions of functional requirements specifications of software to identify their levels of granularity at the early stages of the software development life cycle. To achieve this research objective, the main phases of the research methodology are: • exploratory research: to investigate the impact of the research issue on the approximation of the functional size approximation process; • framework design: to design the framework that assigns scaling factors to functional requirements specifications to identify their levels of granularity; and • framework verification: to verify the usability of the framework by different groups of participants with different experience profiles, and to verify the applicability of the framework with a variety of case studies representing different software systems. The main outcome of this research project is a framework that consists of: a meta-model that identifies the relevant concepts and the relationships that need to be collected by the requirements engineers for achieving full functional specification of software requirements specifications, as well as criteria that identify the levels of granularity of software requirements specifications, and assign scaling factors to rank their levels of granularity. This framework is verified for usability with the same case study by three groups of practitioners in the software engineering industry and verified next for applicability with four case studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.020 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".