The Clinical Utility of the Vulvar Pain Assessment Questionnaire: A Pilot Study
Bibliographic record
Abstract
OBJECTIVE: The aim of the study was to document treatment-seeking experiences of women with chronic vulvar pain, comfort communicating about pain, and test the clinical utility of the screening version of the Vulvar Pain Assessment Questionnaire, screening version (VPAQscreen). MATERIALS AND METHODS: Patients scheduled for an appointment with the Program in Vulvar Health at Oregon Health and Science University were invited to complete the VPAQscreen and answer descriptive questions about previous treatment-seeking experiences and communication with health care providers. Clinicians provided provisional diagnoses based on VPAQscreen summaries, final diagnoses based on gynecological examination, and commented on alignment with clinical observations. Patients gave feedback on the accuracy and helpfulness of the VPAQscreen summary, characteristics of the questions asked, and whether their comfort communicating increased. RESULTS: Participants reported previously seeing approximately 5 medical doctors and 2 other health care providers and perceived them as lacking knowledge of vulvar pain syndromes. Providers indicated that VPAQscreen summaries aligned with clinical presentations and suggested provisional diagnoses with more than 80% accuracy. Participants reported that VPAQscreen summaries were helpful and accurate in summarizing symptoms. Most reported that the number, range, and readability of VPAQscreen questions were good or excellent. More than half reported that completing the VPAQscreen increased comfort when speaking with their Oregon Health and Science University physician. CONCLUSIONS: Patients with vulvar pain often endure a lengthy process of consulting multiple clinicians before securing care. The VPAQscreen was more than 80% accurate in predicting diagnosis at this specialty clinic and was useful in assisting patients with expressing symptoms. The applicability of the VPAQscreen in general practice is unknown, although it shows promise.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".