Reducing search space for Web Service ranking using semantic logs and Semantic FP-Tree based association rule mining
Bibliographic record
Abstract
Ranking and Adaptation (used interchangeably) is often carried out using functional and non-functional information of Web Services. Such approaches are dependent on heavy and rich semantic descriptions as well as unstructured and scattered information about any past interactions between clients and Web Services. Existing approaches are either found to be focusing on semantic modeling and representation only, or using data mining and machine learning based approaches on unstructured and raw data to perform discovery and ranking. We propose a novel approach to allow semantically empowered representation of logs during Web Service execution and then use such logs to perform ranking and adaptation of discovered Web Services. We have found that combining both approaches together into a hybrid approach would enable formal representation of Web Services data which would boost data mining as well as machine learning based solutions to process such data. We have built Semantic FP-Trees based technique to perform association rule learning on functional and non-functional characteristics of Web Services. The process of automated execution of Web Services is improved in two steps, i.e., (1) we provide semantically formalized logs that maintain well-structured and formalized information about past interactions of Services Consumers and Web Services, (2) we perform an extended association rule mining on semantically formalized logs to find out any possible correlations that can used to pre-filter Web Services and reduce search space during the process of automated ranking and adaptation of Web Services. We have conducted comprehensive evaluation to demonstrate the efficiency, effectiveness and usability of our proposed approach.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".