Psychometric Properties of a Peer-Assessment Program to Assess Continuing Competence in Physical Therapy
Bibliographic record
Abstract
BACKGROUND: The College of Physiotherapists of Ontario implemented an Onsite Assessment to evaluate the continuing competence of physical therapists. OBJECTIVE: This study was undertaken to examine the reliability of the various tools used in the Onsite Assessment and to consider the relationship between the final decision and demographic factors. DESIGN: This was a psychometric study. METHODS: Trained peer assessors (n=63) visited randomly selected physical therapists (n=106) in their workplace. Fifty-three physical therapists were examined by 2 assessors simultaneously. The assessment included a review of practice issues, record keeping, billing practices, the physical therapist's professional portfolio, and a chart-stimulated recall process. The Quality Management Committee made the final decision regarding the physical therapist's performance using the assessor's summary report. Generalizability theory was used to examine the interrater reliability of the tools. Correlation coefficients and regression analyses were used to examine the relationships between demographic factors and performance. RESULTS: The majority of the physical therapists (88%) completed the program successfully, 11% required remediation, and 1% required further assessment. The interrater reliability of the components was above .70 for 2 raters' evaluations, with the exception of billing practices. There was no relationship between the final decision and age or years since graduation (r<.05). Limitations Limitations include a small sample and a lack of data on system-related factors that might influence performance. CONCLUSIONS: The vast majority of the physical therapists met the College of Physiotherapists of Ontario's professional standards. Reliability analysis indicated that the number of charts reviewed could be reduced. Strategies to improve the reliability of the various components must take into account feasibility issues related to financial and human resources. Further research to examine factors associated with failure to adhere to professional standards should be considered. These results can provide valuable information to regulatory agencies or managers considering similar continuing competence assessment programs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".