Comparative performance of cobas 4800 HPV Test and Anyplex II HPV HR for high-risk human papillomavirus detection
Bibliographic record
Abstract
ABSTRACT Numerous molecular tests are available to detect human papillomavirus (HPV). We compared the analytical performance of cobas and Anyplex for detection of high-risk (HR) carcinogenic HPV genotypes, assessed the composition of HPV types (other than 16 and 18) that influenced cobas performance, and considered the impact of viral load on test performance. We used data from the Early Detection of Cervical Cancer in Hard-to-Reach Populations of Women Through Portable and Point-of-Care HPV Testing project, which involved collection (2019–2022) of cervicovaginal samples from 1,042 women aged 21–74 years in Belgium ( n = 244), Portugal ( n = 309), Brazil ( n = 244), and Ecuador ( n = 245). Samples were tested by cobas (provides individual results for HPV16 and HPV18 and a pooled result for 12 other HR-HPV types) and Anyplex (provides separate results for 14 HR-HPVs). We calculated HPV positivity by each test and compared performance between tests by calculating Cohen’s kappa statistics. Based on 938 samples with complete data from both tests, positivity rates by cobas were 13.4%, 3.6%, 34.3%, and 45.3% for HPV16, HPV18, 12 pooled HR-HPVs, and any HR-HPV, respectively. Corresponding HPV positivity rates by Anyplex were 14.9%, 3.7%, 37.9%, and 50.0% for the same categories, respectively, with high concordance; kappa statistics were 0.90, 0.87, 0.82, and 0.85, respectively. Based on 355 samples that tested positive for at least 1 of the 12 pooled HR-HPVs, most types showed high agreement (80.9%–100.0%) between individual-Anyplex and pooled-cobas HPV results, except for HPV68 (61.3% agreement). Our findings suggest that the two commercial tests may have different performances, depending on the specific HPV types detected, emphasizing the need for continued research on conditions that may affect these tests, especially for less common or less studied HPV types. IMPORTANCE This study compared two commercial tests—cobas and Anyplex—for detecting high-risk HPV types in women undergoing routine cervical cancer screening or referred for colposcopy. Both tests provide separate results for HPV16 and HPV18, but Anyplex also identifies the remaining 12 high-risk HPV types individually, while cobas groups them together. Overall, we found a high level of agreement between the two tests, supporting their use in clinical practice. However, differences in detecting certain HPV types, particularly those that are less common or less studied, emphasize the importance of choosing the right test. As more countries switch to HPV-based cervical cancer screening, using tests that provide detailed results could help improve risk assessment and optimize patient care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.035 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".