Psychometric properties of the Social knowledge test (SKT) and the Combined stories test (COST) in people with a schizophrenia spectrum disorder
Bibliographic record
Abstract
Background: People with schizophrenia spectrum disorders (SSD) often present with impaired social cognition. Among the measures available to assess these deficits, the Combined stories test (COST) and the Social knowledge test (SKT), that respectively target theory of mind (ToM) and social knowledge, have shown promising psychometric properties in prior studies. Test-retest reliability was however only examined in the general population, and the acceptability of these tests was not previously examined. This study aimed to further document the psychometric properties of the COST and the SKT and the acceptability of these tests in people with SSD and community controls (CO). Methods: Forty-four (44) participants with SSD and 49 CO were administered the COST and SKT twice, about 4 weeks apart, and were asked to rate the acceptability of the tests on a 0 (Very unpleasant) to 10 (Very pleasant) point scale at both timepoints. Results: In both groups, the results revealed an excellent inter-rater reliability and a good test-retest reliability for both tests, though the control non-social reasoning measure included in the COST showed poorer test-retest reliability in the SSD group. Some practice effects were observed but the ToM score from the COST and the SKT total score showed no evidence of ceiling effects at either timepoints. The average acceptability scores ranged between 7.8/10 and 8.3/10 for the COST and between 6.8/10 and 7.9/10 for the SKT. Conclusion: The SKT and the COST present with good psychometric properties, representing good options for future studies or for use in clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.007 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".