MétaCan
Menu
Back to cohort
Record W4411516669 · doi:10.1016/j.scog.2025.100373

Psychometric properties of the Social knowledge test (SKT) and the Combined stories test (COST) in people with a schizophrenia spectrum disorder

2025· article· en· W4411516669 on OpenAlexafffund
Amélie M. Achim, Élisabeth Thibaudeau, Frédéric Haesebaert, Audrey Cayouette, Caroline Cellard

Bibliographic record

VenueSchizophrenia Research Cognition · 2025
Typearticle
Languageen
FieldMedicine
TopicSchizophrenia research and treatment
Canadian institutionsCentre Jeunesse de QuebecBombardier (Canada)Université LavalCentre for Research on Brain Language and Music
FundersFonds de Recherche du Québec - SantéCanadian Institutes of Health ResearchFondation de l’Université Laval
KeywordsSchizophrenia spectrumPsychologyTest (biology)Schizophrenia (object-oriented programming)Classical test theoryPsychometricsClinical psychologyApplied psychologyPsychiatryItem response theoryPsychosis

Abstract

fetched live from OpenAlex

Background: People with schizophrenia spectrum disorders (SSD) often present with impaired social cognition. Among the measures available to assess these deficits, the Combined stories test (COST) and the Social knowledge test (SKT), that respectively target theory of mind (ToM) and social knowledge, have shown promising psychometric properties in prior studies. Test-retest reliability was however only examined in the general population, and the acceptability of these tests was not previously examined. This study aimed to further document the psychometric properties of the COST and the SKT and the acceptability of these tests in people with SSD and community controls (CO). Methods: Forty-four (44) participants with SSD and 49 CO were administered the COST and SKT twice, about 4 weeks apart, and were asked to rate the acceptability of the tests on a 0 (Very unpleasant) to 10 (Very pleasant) point scale at both timepoints. Results: In both groups, the results revealed an excellent inter-rater reliability and a good test-retest reliability for both tests, though the control non-social reasoning measure included in the COST showed poorer test-retest reliability in the SSD group. Some practice effects were observed but the ToM score from the COST and the SKT total score showed no evidence of ceiling effects at either timepoints. The average acceptability scores ranged between 7.8/10 and 8.3/10 for the COST and between 6.8/10 and 7.9/10 for the SKT. Conclusion: The SKT and the COST present with good psychometric properties, representing good options for future studies or for use in clinical practice.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.005
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.409
Threshold uncertainty score0.811

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.005
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0010.007
Science and technology studies0.0010.002
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.034
GPT teacher head0.330
Teacher spread0.296 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueSchizophrenia Research CognitionSame topicSchizophrenia research and treatmentFrench-language works237,207