Supporting the development of critical data literacies in higher education: building blocks for fair data cultures in society
Notice bibliographique
Résumé
In the last ten years digitalized data have permeated our lives in a massive way. Beyond the internet ubiquity and cultural change outlined in what Castells ( 1996 ) called the network society, we are now witnessing a datafied society, where large amounts of digital data—the DNA of information—are driving new social practices. The most enthusiastic discourses on this abundance of data have emphasized the opportunity to generate new business models, with professional landscapes connected to data science and open practices in science and the public space (EMC Education Services 2015 ; Scott 2014 ). However, more recently, the rather naïve logic of data capture and its articulation through various algorithms as drivers of more economical and objective social practices have been the object of criticism and deconstruction (Kitchin 2014 ; Zuboff 2019 ). The university as an institution fell into this paradigm somehow abruptly, while striving to survive its crisis of credibility. The digitalization of processes and services was considered a form of innovation and laid the foundations for the later phenomenon of datafication (Williamson 2018 ). Initially, fervent discourses embraced data-driven practices as an opportunity to improve efficiency, objectivity, transparency and innovation (Daniel 2015 ; Siemens et al. 2013 ). The two main missions in higher education (HE)—teaching and research—went through several processes of digitalization that encompassed data-intensive practices. In teaching, the data about learning and learners collected on unprecedented scales gave rise to educational data mining and particularly to learning analytics (LA) (Siemens and Long 2011 ). While some argued about the value of learning analytics in informing teachers’ decision-making about pedagogical practices as well as learners’ self-regulation (Ferguson 2012 ; Roll and Winne 2015 ), research also uncovered naïve or even poor pedagogical assumptions on the power of algorithms to predict, support and address learning, which were connected to techno-determinist approaches to data (Ferguson 2019 ; Perrotta and Williamson 2018 ; Selwyn 2019 ). The studies in the field have pointed out how few connections there are between LA models and pedagogical theories (Knight et al. 2014 ; Nunn et al. 2016 ), the lack of evaluation in authentic contexts, the scant uptake by teachers and learners (Vuorikari et al. 2016a , b ) and the social and ethical issues connected to the topic (Broughan and Prinsloo 2020 ; Slade and Prinsloo 2013 ; Prinsloo and Slade 2017 ). Moreover, the massive adoption of social media has crossed paths with learning management systems, creating new forms of data of which both teachers and students could be completely unaware (Manca et al. 2016 ).
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,011 | 0,044 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,003 | 0,001 |
| Études des sciences et des technologies | 0,006 | 0,012 |
| Communication savante | 0,018 | 0,013 |
| Science ouverte | 0,003 | 0,005 |
| Intégrité de la recherche | 0,016 | 0,026 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,008 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».