Benefits of Chain of Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medicine Education: An Experimental Evaluation Study with Medical Faculty (Preprint)
Bibliographic record
Abstract
Background: Large language models in artificial intelligence have been among the tools with a significant and real impact on people's daily lives. In this regard, they serve as an aid in specific fields, such as education, helping educators with cumbersome tasks such as periodic evaluations. Objective: This study focused on analyzing the benefits of large language models, particularly the chain-of-thought (CoT) strategy, for the task of evaluating students' Spanish-language medical record writing. The aim was 2-fold: first, we attempted to save time and resources, and second, we used the reasoning of the CoT strategy to evaluate the rubrics and their interpretations. Methods: The proposed solution assessed the application of 2 models-Llama 3.1 and Claude 3.5-in combination with one-shot and CoT to evaluate how medical students write medical records in Spanish. First, machine learning metrics were applied to measure the performance of the solutions. Then, different statistical analyses were performed at the clinical record and item levels. Finally, differences between the proposed models and evaluators were studied in depth. Results: A maximum of 3807 items were evaluated. Claude obtained the best accuracy with slight differences between one-shot and CoT (86.4% and 85.0%, respectively). However, Claude with CoT outperformed the rest of the combinations on all complementary metrics, initially achieving a sensitivity of 94.2%, specificity of 59.5%, precision of 85.8%, and F1-score of 89.6%. Expert review of CoT reasoning determined that 63.8% of the discrepancies were model hits, raising Claude's final accuracy to 94.6% (SD 4.3%). In the final phase, sensitivity was 98.0% (SD 2.3%), specificity improved to 83.3% (SD 14.3%), and F1-score reached 96.2% (SD 3.3%). Sectional analysis showed greater difficulties in the "History of present illness" section (n=125 discordances). Conclusions: CoT demonstrated strong potential for supporting the evaluation of clinical records written in Spanish by medical students and providing feedback to them. More importantly, it showed significant promise in assisting professors by assessing the quality of their rubrics and identifying possible errors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.028 | 0.206 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".