A study examining inter-rater and intrarater reliability of a novel instrument for assessment of psoriasis: the Copenhagen Psoriasis Severity Index
Bibliographic record
Abstract
BACKGROUND: There is a perceived need for a better method for clinical assessment of the severity of psoriasis vulgaris. The most frequently used system is the Psoriasis Area and Severity Index (PASI), which has significant disadvantages, including the requirement for assessment of the percentage of skin affected, an inability to separate milder cases, and a lack of linearity. The Copenhagen Psoriasis Severity Index (CoPSI) is a novel approach which comprises assessment of three signs: erythema, plaque thickness and scaling, each on a four-point scale (0, none; 1, mild; 2, moderate; 3, severe), at each of 10 sites: face, scalp, upper limbs (excluding hands and wrists), hands and wrists, chest and abdomen, back, buttocks and sacral area, genitalia, lower limbs (excluding feet and ankles), feet and ankles. OBJECTIVES: To evaluate the inter-rater and intrarater reliability of the CoPSI and to provide comparative data from the PASI and a Physician's Global Assessment (PGA) used in recent clinical trials on psoriasis vulgaris. METHODS: On the day before the study, 14 dermatologists (raters) with an interest in psoriasis participated in a detailed training session and discussion (2.5 h) on use of the scales. On the study day, each rater evaluated 16 adults with chronic plaque psoriasis in the morning and again in the afternoon. Raters were randomly assigned to assess subjects using the scales in a specific sequence, either PGA, CoPSI, PASI or PGA, PASI, CoPSI. Each rater used one sequence in the morning and the other in the afternoon. The primary endpoint was the inter-rater and intrarater reliability as determined by intraclass correlation coefficients (ICCs). RESULTS: All three scales demonstrated 'substantial' (a priori defined as ICC > 80%) intrarater reliability. The inter-rater reliability for each of the CoPSI and PASI was also 'substantial' and for the PGA was 'moderate' (ICC 61%). The CoPSI was better at distinguishing between milder cases. CONCLUSIONS: The CoPSI and the PASI both provided reproducible psoriasis severity assessments. In terms of both intrarater and inter-rater reliability values, the CoPSI and the PASI are superior to the PGA. The CoPSI may overcome several of the problems associated with the PASI. In particular, the CoPSI avoids the need to estimate a percentage of skin involved, is able to separate milder cases where the PASI lacks sensitivity, and is also more linear and simpler. The CoPSI also incorporates more meaningful weighting of different anatomical areas.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".