Affordances for expressing collective identities - The case of the 2024 French parliamentary elections
Notice bibliographique
Résumé
In this project we studied how users express collective identities on social media, in particular in the context of the recent French legislative election (June-July 2024). Our main data set was from X (see section 2). In some of the subprojects we also looked at Tiktok (subprojects 1,3,5) and Instagram (subproject 1).The project is part of the HORIZON Europe project “Social Media for Democracy (SoMe4Dem) – understanding the causal mechanisms of digital citizenship”. The SoMe4Dem project studies the impact of social media along three dimensions: participation, polarisation and trust by developing democracy theory, providing evidence for social media impact, performing experiments and building models to reveal causal mechanisms, and studying practices to improve digital citizenship.As part of this project we want to understand how the technical and social affordances of social media platforms relate to the different functions of the public sphere in liberal democracies, such as information, deliberation and the forming of collective actors. The formation and expression of collective identities is part of the latter and the focus of this project.Starting point of the analysis was the identification of political communities in the retweet network. Beginning with the seminal work by Conover et al. (2011) it was shown that people are more likely to share content the more the content corresponds to their own beliefs. Therefore, in political debates, clusters in the retweet network can be often interpreted as clusters of accounts with a similar political stance. Conover et al. (2011) showed this for Democrats and Republicans in the US context, Gaumont et al. (2018) for the French presidential elections 2017 and Gaisbauer et al. (2021) for the Saxon state elections in Germany 2019. In a next step we compared between the communities how people use different affordances for identity expression such as user biographies, emojis in screen names, or profile pictures.More specific aspects were studies in several subprojects: One project studies how climate activists and feminists express their identities and how feminism and climate change is addressed in the different communities, a second sub-project asked to which extend users express political identities. A third subproject to which extent and how users express use geographical information in their bios, for instance to express regional identities and whether there are systematic differences between rural and urban areas. The fourth subproject presented a case study comparing the self-presentation of a journalist between X, Instagram and TikTok. In the fifth subproject we use a seeded Structural Topic Model to map differences in Bio self-description patterns reflecting the political identities of two major communities ("Rassemblement National" and "Socialistes”) previously identified through the retweet network.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,013 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,001 | 0,002 |
| Études des sciences et des technologies | 0,005 | 0,004 |
| Communication savante | 0,006 | 0,004 |
| Science ouverte | 0,001 | 0,003 |
| Intégrité de la recherche | 0,002 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,008 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».