Reliability and validity of the instrument for scoring clinical outcomes of research for epidermolysis bullosa (iscorEB)
Bibliographic record
Abstract
BACKGROUND: Epidermolysis bullosa (EB) is a group of rare and currently incurable genetic blistering disorders. As more pathogenic-driven therapies are being developed, there is an important need for EB-specific validated outcomes measures designed for use in clinical trials. OBJECTIVES: To test the reliability and construct validity of an instrument for scoring clinical outcomes of research for EB (iscorEB), a new combined clinician- and patient-reported outcomes tool. METHODS: We conducted an observational study consisting of independent 1-day assessments (six assessors) at two academic hospitals. The assessments consisted of iscorEB clinician (iscorEB-c), Birmingham Epidermolysis Bullosa Severity (BEBS) and global severity assessment for physicians; and iscorEB patient (iscorEB-p), Quality of Life evaluation in Epidermolysis Bullosa and Children's Dermatology Life Quality Index for patients. Construct validity and intraclass correlation coefficients (ICCs) for interobserver, intraobserver and test-retest reliability were calculated. RESULTS: Overall, 31 patients with a mean age of 19·5 years (1·8-45·2) were included. Disease severity was mild in 42% of cases, moderate in 29% and severe in 29%. The interobserver ICC was 0·96 for both the clinician-reported section of iscorEB-c and BEBS. The ICC for intraobserver reliability was 0·91 and 0·70 for the skin and mucosal domains of iscorEB-c, respectively. Cronbach's alpha for iscorEB-c was 0·89. The test-retest reliability of iscorEB-p was 0·97 and Cronbach's alpha was 0·84. The clinical score differentiated between subjects with mild, moderate and severe disease, and both clinical and patient subscores discriminated between recessive dystrophic EB and other EB subtypes. CONCLUSIONS: iscorEB has robust reliability and construct validity, including strong ability to distinguish EB types and severities. Further studies are planned to test its responsiveness to change.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".