Social validation in aphasiology: Does judges' knowledge of aphasiology matter?
Bibliographic record
Abstract
Background: Social validity assessments can be used to examine clinical significance of changes due to treatment of aphasia. Behavioural researchers have noted the need to investigate various methodological issues in social validity research. For example, differences in rater characteristics have been noted to influence social validation ratings of treatment outcomes.Aims: This study examined the possibility of differences in social validity ratings across judges with varying degrees of knowledge of and experience with aphasia. The secondary purpose was to replicate results of previous research that showed the clinical significance of communication partner training. Research questions included: (1) Does the level of knowledge of aphasia and experience with persons with aphasia result in a difference in social validity ratings of pre- and post-training conversations between a student volunteer and an elder with aphasia? (2) Will the significant social validity findings previously obtained from members of the extended community be replicated (Hickey, 2000)?Methods & Procedures: Ten naïve individuals (no familiarity with aphasia), ten second-year graduate students majoring in speech-language pathology (some familiarity with aphasia), and ten Speech-Language Pathologists served as judges. After watching two pre- and two post-training videotaped conversations, the judges provided ratings for seven dimensions of conversations to examine clinical significance of changes in pre- and post-training conversations between a student volunteer and an elder with aphasia. A mixed design with between and within subjects effects, and interaction effects was used.Outcomes & Results: Repeated measures ANOVA revealed significant main effects for group on two items, significant main effects for training on all seven items, and significant interaction effects for five items. Pre-training ratings showed greater variability than post-training ratings. Naïve judges provided the lowest pre-training ratings, and generally, the most change in pre- and post-training ratings. Post-training ratings of the three groups became more similar.Conclusions: This study suggests that truly naïve judges who are representative of the general public may provide the most robust findings in social validation studies of aphasia treatment outcomes. However, further research is needed to determine the source of variability beyond level of knowledge of aphasia. This study also replicated the results of previous research by revealing the clinical significance of communication partner training for elders with aphasia.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.093 | 0.318 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".