The effect of rater training on scoring performance and scale-specific expertise amongst occupational therapists participating in a multicentre study: a single-group pre–post-test study
Bibliographic record
Abstract
PURPOSE: In order to enhance the quality of the data collected in a multicentre validation study of a revised Danish version of the McGill Ingestive Skills Assessment (MISA), the authors developed a rater training programme. The purpose of the present study was to evaluate the effect of the training on scoring performance and scale-specific expertise amongst raters. METHOD: During 2 days of rater training, 81 occupational therapists (OTs) were qualified to observe and score dysphagic clients' mealtime performance according to the criteria of 36 MISA-items. The training effects were evaluated pre- to post-training using percentage exact agreement (PA) of scored MISA items of a case-vignette and a Likert scale self-report of scale-specific expertise. RESULTS: PA increased significantly from pre- to post-training (Z = -4.404, p < 0.001), although items for which the case-vignette reflected deficient mealtime performance appeared most difficult to score. The OTs scale-specific expertise improved significantly (knowledge: Z = -7.857, p < 0.001 and confidence: Z = -7.838, p < 0.001). CONCLUSION: Rater training improved OTs scoring performance when using the Danish MISA as well as their perceived scale-specific expertise. Future rater training should emphasis the items identified as those most difficult to score. Additionally, further studies addressing different training approaches and durations are warranted. IMPLICATIONS FOR REHABILITATION: When occupational therapists (OTs) use the McGill Ingestive Skills Assessment (MISA) they observe, interpret and record occupational performance of dysphagic clients participating in a meal. This is a highly complex task, which might introduce unwanted variability in measurement scores. A 2-day rater training programme was developed and this builds on the findings of several studies. These suggest that combinations of different training methods tend to yield the most effective results. Participation in the newly developed training programme on how to administer the MISA significantly reduces unwanted variability in measurement scores and improves OTs' competency. The training programme could be used in undergraduate and postgraduate dysphagia education initiatives to help OTs understanding of the content and the scoring criteria for each aspect of occupational performance during a meal, thus developing observation skills as well as recognizing and avoiding the most common errors in measurement scores.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".