SCYLLA OR CHARYBDIS: NAVIGATING BETWEEN EXCESSIVE EXAMINATION AND NAÏVE RELIANCE ON SELF‐ASSESSMENT
Bibliographic record
Abstract
A Peculiar Disjunction Is Becoming Apparent In Discussions About Assessment Of Competence In The Health Professions. On One Hand, Over The Last 30 Years, There Has Been An Explosion Of Testing Technologies Such That Health Professionals Undergo An Almost Endless Series Of Examinations During Training And Practice. Written Tests, Such As The Ubiquitous Multiple-choice Examination, And Performance Tests, Such As The Objective Structured Clinical Examination, Are Used For Admissions, At Regular Intervals In Training, For Certification, Licensure And, In Some Professions, For Maintenance Of Certification. Several Objections Can Be Raised About This ‘culture Of Examination’, Including The Distorting Effect Of Examination On Student Learning; The Observation That Behavioural Checklists Penalize Experts Who Use Pattern Recognition And Synthesis; The Virtual Disappearance Of Feedback As Examination Banking And Testing Security Are Employed; And, Finally, The Contracting Out Of High-stakes Examinations And Examination Preparation, Resulting In Enormous Increases In Costs For Individual Health Professionals. Health Professionals Today Live In What Michel Foucault (1975/1995) Called An ‘examined Society’ In Which Constant Surveillance And Testing Locates The Responsibility For Competence Externally To The Individual. Disturbingly, Examination Developers Often Ignore Important Research About These Limitations And Side-effects Of Examination Technology (Norman 2005; Schuwirth And Van Der Vleuten 2006). Simultaneously There Is A Very Different Discourse About Assessment That Is Tethered To A ‘trinity’ Of Reflective Technologies: Self-assessment, Self-direction And Self-regulation (Hodges 2004). Emerging From Adult Learning Theory, This Discourse Is Centred On The Idea That The Locus For Control Of Competence Is Internal (Norman 1999). This Model Has Spawned A Very Different Set Of Assessment Technologies That Includes Portfolios, Reflective Diaries, Logbooks, Self-directed Web-based Modules, Etc. Oddly, Enthusiasts For This Approach Seem To Give Little Consideration To The Enormous Literature Showing That Self-assessment Is Actually Very Poor. In Many Studies, A Large Number Of Learners Can Be Found Who Appear Unable To Identify Their Own Strengths And Weaknesses (Davis Et Al. 2006). Furthermore, Even When Self-assessment Is Possible, It Does Not Necessarily Lead To Improvements In Practice (Davis Et Al. 1995). Eva And Regehr (2005) Have Called For A Reconceptualization Of Self-assessment And Nelson And Purkis (2004) Have Argued The Folly Of Basing Systems Of Competence Assessment On The Tenuous Process Of ‘reflection’. While These Discursive Trains Appear To Be Headed Down Completely Different Tracks, We Are Confusing Students. High-stakes, Summative Assessments Are No Doubt Essential To Discover And Remove A Few Very Low Performers (Although The Eventual Fail Rate On Most Health Professional Examinations Is Near Zero); To Reassure The Public That The Professions Are Seriously Monitoring The Knowledge And Skills Of The Next Generation; And For Purposes Of Standard Setting Across The Country And Internationally. On The Other Hand, The Testing Machine That Has Been Created Has Almost Nothing To Do With Learning. In The Absence Of Meaningful Feedback, And With Levels Of Test Security That Completely Disconnect Examinations Both Spatially And Temporally From Where Learning Occurs, Big Exams Contribute Very Little To Improving Individual Competence. Yet Living In The Comfortable Delusion That Individual Health Professionals Will Engage In A Continuous Process Of Self-reflection, Identify Their Strengths And Weakness, Find Self-directed Learning Opportunities To Remedy Them And Thereby Improve Their Practice Leaves The Reader Of Literature On Self-assessment Incredulous, If Not Horrified. How Might We Imagine A Rapprochement Of These Two Poles? How Does The Health Professional Educator Navigate Between The Scylla Of Excessive External Examination And The Charybdis Of Naïve Reliance On Self-assessment? It May Be That The Solution Lies In Some Form Of Guided Self-assessment — Just Enough External Input To Correct For The Vagaries And Inaccuracies Of Self-judgement, But Not So Much That It Deforms The Exercise To Meet The Needs Of Examiners Or Institutions. Working With A Mentor, Teacher Or Coach Who Can Help A Student (Or A Professional In Practice) Be Sure They Know ‘when To Slow Down’ (Moulton Et Al. 2007) Or ‘when To Look It Up’ (Eva And Regehr 2007) May Be The Basis Of A More Evidence-based Approach To Competence Assessment. Notwithstanding The Potential For Abuses Of Power When One's Self-reflections Are ‘guided’ By Someone Else (Hodges 2004), This Approach Would Seem To Fit Better With Evidence Of How Competence Works.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.035 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".