Expérimentation d’un modèle d’évaluation certificative dans un contexte d’enseignement scientifique
Bibliographic record
Abstract
Le programme quebecois de science et technologie est base sur une approche par competences. Ce choix implique des defis importants et principalement quand l’evaluation est de nature certificative. Une des competences a evaluer concerne l’investigation scienti ‐ fique. En s’appuyant sur les travaux de Rey et al . (2003), nous avons concu un modele d’evaluation qui permet de juger du developpement de cette competence. Differentes situa ‐ tions d’evaluation ont ete crees et administrees aupres de 560 eleves du secondaire pour verifier si le modele : (1) est adequat pour mesurer le niveau de competence des eleves et (2) se comporte de la meme facon selon le contexte disciplinaire. Les resultats montrent que le modele permet de classer les eleves selon trois niveaux de maitrise : competence assuree, competence partielle et maitrise des habiletes. Mots cles : evaluation, competences, investigation scientifique The Science and Technology curriculum in the Province of Quebec, based on competencies, represents a challenge for science teachers, particularly for high ‐ stakes assessment. Teach ‐ ers who do know how to conduct hands ‐ on assessment must deal with practical con ‐ straints. In this context, we adapted Rey’s et al. (2003) work to construct an assessment model in relation with the Scientific Inquiry Competence. We designed different assess ‐ ment situations which we administered to 560 junior high school students to verify whether a) the model is helpful in assessing learners’ level of competency in scientific in ‐ quiry b) the results are comparable among disciplines on which the assessment situations are based. The results show that the model works as predicted for different learners’ levels: Full competency, Partial Competency, Skill. Key words: assessment, competencies, scientific inquiry
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".