Thematic Working Group 5: Formative assessment supported by technology.
Bibliographic record
Abstract
The future of assessment faces major challenges including the use of IT to facilitate \nformative assessment that is important for improving learners’ development, motivation \nand engagement in learning. In many countries, in recent years, a renewed focus on \nassessments to support learning has been pushing against the burgeoning of testing for \naccountability, which in some countries, renders effective formative assessment \npractices almost impossible. Moreover, a systematic review by Harlen and Deakin Crick \n(2002) revealed that a strong focus on summative assessment for accountability can \nreduce motivation and disengage many learners. At the same time use of IT‐enabled \nassessments has been increasing rapidly, as they offer promise of cheaper ways of \ndelivering and marking assessments as well as access to vast amounts of assessment \ndata from which a wide range of judgements might be made about students, teachers, \nschools and education systems (Gibson & Webb, 2015). These opportunities also extend \nto assessment of complex collaborative work (Webb & Gibson, 2015). Current \nopportunities for using IT, including for harnessing the data that is being collected \nautomatically, for formative assessment are underexplored and less well understood \nthan those for summative assessments. Opportunities for learning with IT and perhaps \nwith less teacher input are increasing but this depends on students developing as \nautonomous or independent learners. Research in formative assessment including \neffective feedback has emphasised the value of peer assessment practices for \ndeveloping self‐assessment capabilities and hence independent learners (Black, \nHarrison, Lee, Marshall, & William, 2003). At previous EDUsummITs the possibilities and \nchallenges for IT‐enabled assessments to support simultaneously both formative and \nsummative purposes were analysed (Webb, Gibson, & Forkosh‐Baruch, 2013). While these challenges remain, at EDUsummIT 2017 we focused on the opportunities and \nchallenges of IT supporting formative assessment because effective formative \nassessment is known to be extremely important for learning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.006 |
| Science and technology studies | 0.009 | 0.006 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".