Notice bibliographique
Résumé
As members of a learned profession, engineers are often required to assess and critique the work of others. Preparation for this professional responsibility should be developed during their academictraining, alongside other required skills. This authorproposes that there are generic skills and trainingmethodologies that can be applied to both technical and“soft skill” situations to prepare students for this task.This paper discusses results of a peer assessment exerciseapplied to a “soft-skills” situation.The main objectives of this experiment were to i)develop peer assessment skills in students, ii) maintain orimprove the accuracy of assessments for subjectivematerial, iii) improve students’ skills in the subject area,and iv) potentially reduce marking effort for instructors.The experiment described in this paper involved peerassessment of a short report (3 – 5 pages) required as aterm assignment in a senior course on ethics andprofessionalism. The reports were prepared andsubmitted by groups of two students. Each student wasthen randomly assigned two other reports to assess in adouble-blind fashion, except that no student reviewerreceived their own report. For reference and analysis,each report was also assessed by both the instructor anda Teaching Assistant resulting in approximately sixseparate assessments per report The results were used todetermine a grade for the assignment. The originalassignment rubric was used for all assessments. Inaddition, formative feedback was provided by thereviewers and returned to the authors.The quality of the numerical results was analyzed bycomparing the marks determined by the student assessorsto the reference (instructor, TA) assessments. An averagedifference of 8.5% was observed, and was consideredgenerally acceptable given the subjective nature of thematerial. Student “generosity bias” was also considered,but found to be virtually non-existent with a difference instudent versus reference averages of less than 0.2%.“Outliers” were anticipated, and student assessmentshowed approximately twice the standard deviation of thereference marks. A weighted average was used todetermine the assignment mark, and any marks outside a20.0% band were de-weighted. Approximately 25% ofcases were weight-adjusted, resulting in a maximum markadjustment of 4.1% and an average adjustment of only1.6%.Feedback was solicited from students prior to the peerreview period and at the end of term. Informal feedbackwas solicited prior to the review period regardinginstructions and logistics, and was used to refine the setupfor the peer review phase. Questions on the value of boththe exercise and the feedback provided were included inan end-of-term survey of students about the course, with83% finding the exercise “a bit” or “quite” educationaland 74% finding the peer feedback “a bit” or “quite”helpful.Involving students in this peer evaluation exercise hadgenerally positive outcomes and provided experiencefrom which to improve future implementation of peerassessments to achieve the objectives of this experiment.Recommendations regarding future application include:importance of instructions and setup, student training and rehearsal, and mark determination considerations.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».