Comprehension and construct validity of the Visual Prostate Symptom Score by men with obstructive lower urinary tract symptoms in rural Africa
Bibliographic record
Abstract
INTRODUCTION: The Visual Prostate Symptom Score (VPSS) is an image-based interpretation of the International Prostate Symptom Score (IPSS) intended to quantify frequency, nocturia, weak stream, and quality of life (QoL) in a literacy-independent manner. METHODS: Ugandan men presenting with lower urinary tract symptons (LUTS) to a rural clinic completed VPSS and IPSS independently and then with assistance. They verbally interpreted VPSS images, rated question usefulness, and suggested improvements. Responses between word-based and image-based measures were compared (Student's T, Fisher's exact, and Spearman's correlation tests). RESULTS: One hundred thirty-two scores from 33 men (mean age: 61 years, range 28-93; education: no schooling 20%, grades 1-4 62%, 5-7 9%, 8-12 9%). Correlation between IPSS and VPSS scores was positive (r=0.70), as was that between the individual irritative, obstructive, and QoL questions. Independent of education, the weak stream image was best-recognized. Likert scale measures indicated this was the most useful image, followed by daytime frequency. Nocturia and QoL images were rated as less clear, with explanation required before most understood that QoL facial expression images reflected overall LUTS impact. Improvements suggested included: increased image size for frequency and nocturia pictograms, increased black/white contrast for nocturia, and addition of an image to allow reporting of urgency. CONCLUSIONS: In a population with little formal education, there was positive correlation between IPSS and VPSS, with inherent recognition best for weak stream and worst for QoL images. Increased image clarity and an additional image for urgency will enhance the global utility of the VPSS for men to report symptoms of LUTS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".