A Composite Cytology–Histology Endpoint Allows a More Accurate Estimate of Anal High-Grade Squamous Intraepithelial Lesion Prevalence
Bibliographic record
Abstract
BACKGROUND: There is debate about the accuracy of anal cytology and high-resolution anoscopy (HRA), in the diagnosis of anal human papillomavirus (HPV)-related squamous intraepithelial lesions (SIL). Few studies have performed both simultaneously in a large sample of high-risk individuals. METHODS: At baseline in a community-based cohort of HIV-infected and uninfected homosexual men ages ≥35 years in Sydney, Australia, all men underwent anal swabbing for cytology and HPV genotyping, and HRA-guided biopsy. We evaluated the separate and combined diagnostic accuracy of cytology and histology, based on a comparison with the prevalence of HPV16 and other high-risk (HR) HPV. We examined trends in HPV prevalence across cytology-histology combinations. RESULTS: Anal swab, HRA, and HPV genotyping results were available for 605 of 617 participants. The prevalence of cytologically predicted high-grade SIL (HSIL, 17.9%) was lower than histologically diagnosed HSIL (31.7%, P < 0.001). The prevalence of composite-HSIL (detected by either method) was 37.7%. HPV16 prevalence was similar in men with HSIL by cytology (59.3%), HSIL by histology (51.0%), and composite-HSIL (50.0%). HPV16 prevalence was 31.1% in men with composite-atypical squamous cells suggestive of HSIL, to 18.5% in men with composite-low-grade SIL, to 12.1% in men with composite-negative results (Ptrend < 0.001). CONCLUSIONS: Significantly more HSIL was detected when a composite cytology-histology endpoint was used. Increasing grade of composite endpoint was associated with increasing HPV16 prevalence. IMPACT: These data suggest that a composite cytology-histology endpoint reflects meaningful disease categories and is likely to be an important biomarker in anal cancer prevention. Cancer Epidemiol Biomarkers Prev; 25(7); 1134-43. ©2016 AACR.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".