From Scenario to Code: Structured Prompting for LLM-Based Unit Test Generation
Bibliographic record
Abstract
Automated unit test generation is crucial for improving software quality. Existing tools such as EvoSuite exhibits several limitations by generating tests that may lack readability and clarity. These tools may also struggle to cover specific code branches, particularly when complex code components are targetted. Furthermore, studies have shown that the generated unit tests may contain incorrect assertions or cause unexpected behaviors. To address these challenges, we propose a novel approach referred as Two-Step Zero-Shot Prompting (2SZSP), which leverages large language models (LLMs) to structure test generation by identifying relevant test scenarios and generating the corresponding unit test code. Our approach was evaluated on multiple object-oriented projects written in Java from the SBST 2020 dataset and compared to EvoSuite. The results show that, while our approach achieves lower code coverage, it generates more readable, better-contextualized tests with higher mutation score suggesting an increased ability to detect faults. These results pave the way for the integration of LLMs into automated testing pipelines, with the potential to improve the relevance of generated tests and their impact on software robustness.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.016 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".