In the service of the stakeholder: a critical, mixed-method program of research in high-stakes language assessment
Notice bibliographique
Résumé
The three studies presented here represent a two-year program of research that critically explored one case of high-stakes language assessmentâEnglish proficiency assessment for teacher certification in Quebec, Canada.The first study was an examination of the final administration of a writing test used for this purpose at one Quebec university before this test was replaced. In this study, mismatches were revealed in stakeholder perceptions of the task to be produced in this assessment. The second study examined the pilot administration of a new replacement test. It focused on the socio-political environment of the testânamely, how the perception of high or low stakes by raters affected scoring. Results from this study suggested that test stakesâto all stakeholders, including ratersâis a worthwhile focus of study. The third study examined the first official administration of the new test, and focused on rater behavior from a socio-cognitive perspective, suggesting that information on decision-making style may provide insight into variability in rater scoring.This program of research has been critical in that â¢it has integrated social and political values into the test validation process;â¢test stakeholders, including the test takers themselves, have not only been consulted but have determined the direction of the research program to a great extent; andâ¢the conflicting views and competing interests of the stakeholders have been embraced and have enriched the research program.These studies will make contributions to the field of language assessment, and in particular, in better understanding how all elements of the subjectively-scored assessment situation interact. Because of its critical approach, these studies have demonstrated a responsible and progressive approach to researching the assessment of language proficiency for professional certification. In addition, the use of mixed methods designs in all three studies has been somewhat innovative and will add to the emerging field of mixed method research.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,011 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,001 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,003 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».