Bibliographic record
Abstract
UML's support for modeling inter-use-case semantics is very limited.Briand and Labiche (2002) propose the TOTEM methodology to address the generation of tests for such dependencies.However, their generation algorithm, which heavily relies the notion of interleaving, presents one major problem: even for small examples, the number of sequences of use cases to test is unmanageable.The question we ask is: is it feasible to capture use cases and their sequential dependencies in a specification whose executability can allow system-level testing while avoiding the explosion of the number of use case sequences to generate?To answer positively this question we develop a specification for TOTEM's library system in Arnold and Corriveau's ACL specification language and explain how this specification can be used to generate a manageable number of tests.The key idea is that by modeling use case mutual independence we can considerably reduce the number of tests.For example, Robert Binder (2000) has introduced the notion of extended use cases (eUCs) for system-level testing.Roughly put, an eUC can be thought of as a parameterized use case, the parameters of each use case corresponding to the path sensitization variables required to cover the different scenarios (i.e., paths) through that use case (Ibid.).In UML, the requirements of a system are to be modeled using a set of use cases.These use cases, as well as their relationships (i.e., between themselves and with actors) are captured in a use case diagram (Pilone, 2013).Unfortunately, it is widely acknowledged that UML's support for modeling inter-use-case semantics is very limited.In fact, it consists of only two built-in stereotypes: i) a use case may include another one and ii) a use case may extend another one (Ibid.).The exact semantics of these relationships is still debated.More importantly, these two stereotypes are not sufficient to address (the possibly complex) sequential dependencies between use cases, as explained at length by Ryser and Glinz (2000).Binder (2000) defines system testing as concerned with testing an entire system based on its specifications.This task involves the validation of both functional and non-functional (e.g., performance) requirements.In their seminal paper on system testing using UML, Briand and Labiche (2002) remark:[T]hough system testing techniques are in principle implementation-independent, they depend on the notations used to represent the system specifications.In the context of UML-based object-oriented analysis, it is then necessary to develop techniques to derive system test requirements from analysis models such as use case models, interaction diagrams, or class diagrams.To this end, these authors propose the TOTEM methodology that addresses the generation of test requirements from a set of UML artifacts namely: a use case diagram, use case descriptions (in natural language), a sequence diagram for each use case, class diagrams,
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.022 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.006 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".