Bibliographic record
Abstract
There are a number of different approaches to runtime validation, and one such method is Model Based Testing (MBT), a quality assurance technique where a test suite is generated from an abstract model.While there exists a number of different approaches to accomplish modelbased testing, most are state-based (and thus faced with problems such as state explosion, correspondence to code, etc.).In this thesis, we instead focus on an alternative approach to MBT, namely scenario based testing, which we consider in the context of runtime validation.More specifically, we address scenario monitoring and validation in ACL/VF ACL/VF.This approach, developed by Dr. Corriveau and his research team, provides both a language (ACL -Another Contract Language) to specify an implementation-independent testable model and a tool (VF -the Validation Framework) to validate an implementation against an ACL model.The two current versions of the ACL/VF have a number of issues that prevent it from being a usable solution.The original version is .NET specific and incorporates external tools that are no longer supported.Upgrading it to a more recent version of .NET would amount to a complete rewrite.But, more importantly, this initial version has a major bug in its way of monitoring scenarios and a solution requires rethinking a significant and highly technical portion of the implementation of the VF.Instead a new version was implemented and tested with the JavaMOP framework.That solution requires manually mapping an ACL specification to a corresponding set of JavaMOP monitors.Experimentation with this second solution however revealed it cannot manage and monitor multiple scenarios running simultaneously!In light of these very significant difficulties, the most immediate research question for ACL/VF is to determine whether it (and more specifically its rich scenario monitoring) can be implemented in a commonly-used environment such as the core Java/J2EE framework (ideally with no dependencies on external tools).This thesis provides an affirmative answer to this question: a case study is used to illustrate the proposed approach to monitor ACL scenarios using Java threads.We also provide a one to one mapping of ACL elements to corresponding Java/J2EE based code.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".