Bibliographic record
Abstract
Objective: Self-assessment is a readily available evaluation method that is seldom used in postgraduate education because of its poor concurrent and predictive validity. Self-assessment tools often use global performance checklists rather than specific statements of expected competencies. Our goal was to develop and implement a self-assessment tool for residents based on rotation-specific learning objectives. Description: In 1998 at Queen's University, a set of self-assessment tools for core surgical residents was developed, each based on the learning objectives of one specific clinical rotation. Before and after each core surgical rotation, residents estimated their competencies on the objectives of that rotation using a scale of 1 (totally incompetent) to 7 (totally competent). Immediately after each rotation, trainees also scored their perceived competencies on a set of 15 generic (rotation-independent) core surgical objectives. The assessment sheets were collected immediately after the in-training evaluation but were not used in the formal evaluation process. Discussion: In the first year of implementation, nine core residents rotating through seven surgical services completed 43 self-assessments. As expected, the second-year residents' mean scores were higher than the first-year residents' scores, and the mean post-rotation scores were higher than the prerotation scores. There was, however, significant variability among services, and statistically significant improvements were found for only two of the seven services. Self-reported gains in competency can provide valuable information for specific rotation evaluation and feedback. One expects to see a gradual improvement in the generic objectives through the two core training years. Trainees whose self-assessment appears to lag or decline will prompt focused discussion among program coordinators or advisers to identify the underlying problem(s). Self-assessment of individual rotation objectives provides significant feedback both to trainees and to faculty. Trainees readily identify their areas of weakness, and when such assessments are used formatively mid-rotation, they can deliberately focus their efforts on the problem areas. Teaching faculty can easily identify objectives that are consistently covered inadequately on their services, a process that can lead to exploration of underlying issues and potential program enhancements. Self-assessment can form an excellent basis for constructive, bidirectional, feedback. It allows faculty and trainees to explore together both true competencies as well as individual—and possibly erroneous—perceptions of competency. On their part, trainees can use their assessments to identify service-related issues worth discussing. Evaluation: The residents found the forms easy to complete, and the initial response rate was satisfactory. Compliance was far from ideal, however, and methods to enhance it are being sought. The system was easy to implement and required only minimal resources. The key issue that needs to be addressed next is that of concurrent validity of the tool with in-training evaluations and objective scores; the exact role that this self-assessment tool can play within the post-graduate evaluation process will then be possible to define.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.018 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".