How AI Changes Online Knowledge Sharing
Notice bibliographique
Résumé
Generative AI (GenAI) presents remarkable capabilities in tasks involving knowledge work, particularly its ability to create content and assimilate large volumes of structured and unstructured data. This technology is therefore reshaping how people create, retrieve, share, and apply knowledge (Benbya, Strich & Tamm, 2024). One particularly affected area of knowledge work by GenAI is that of online communities (OCs). OCs are home for users who join to share knowledge, support each other, collaborate, and innovate (Faraj, Jarvenpaa & Majchrzak, 2011). I seek to develop a research agenda on the effect of GenAI on the dynamics of voluntary knowledge sharing online. To do that, I propose three roles that GenAI plays in OCs, and then I invite scholars to think of pertinent questions we can ask about these roles. This stream of research carries significant importance for knowledge sharing and AI development. Three characteristics make OCs uniquely affected by GenAI: their public nature, the large volume of the knowledge they produce, and the heavy reliance on interaction for survival. The publicly available large volume of knowledge makes OCs fertile grounds for training AI models, which brings benefits in knowledge aggregation and sharing, but also raises concerns about intellectual property rights and the opportunism of AI corporations that gain profit from the effort of volunteers. The heavy reliance on interaction makes the social fabric of OCs sensitive to the injection of AI agents as alternative interaction partners, thus possibly improving or destroying the social structure depending on whether AI agents support or replace peers. The three characteristics of OCs help us think of GenAI’s role and the influence it can have on them. Namely, three roles are pertinent: AI as a shaper, substitute, or inhibitor of knowledge sharing. As a shaper, GenAI influences how OC knowledge is filtered, categorized, and summarized for users. As a substitute, users seek the GenAI agent as an exchange partner, not only as a tool to access the OC. As an inhibitor, GenAI discourages users from participating publicly. The three GenAI roles create benefits and dangers for continuing knowledge sharing and motivate new research questions. As a shaper, GenAI can protect users from cognitive load by facilitating the search for pertinent discussions and experts and providing personalized experiences. However, it can create bias towards popular information and limit diversity and innovation. Research is needed to understand how users’ participatory behavior changes with AI. Next, as a substitute, GenAI can help streamline simple or redundant questions, fill gaps when experts are missing, or fact-check. However, over-reliance on it can erode social dynamics. Research is needed to understand why and when AI agents may be preferable interaction partners, and when they would be beneficial or detrimental to knowledge work. Finally, the inhibitor AI can detect inappropriate behavior and deter detrimental participation. However, mistrust in AI tools that feed on volunteer efforts can lead users to self-censor. Research is needed to investigate the appropriate governance models of OCs in the age of AI, especially when dealing with intellectual property and fairness towards the OC. Studying this topic is crucial for knowledge sharing and AI development. It allows us to integrate AI into knowledge-intensive work responsibly. Additionally, as volunteer-built repositories are the original fuel for AI training. We need to think how the influence of AI use would change how future models learn.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,039 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,003 | 0,003 |
| Études des sciences et des technologies | 0,006 | 0,006 |
| Communication savante | 0,016 | 0,025 |
| Science ouverte | 0,002 | 0,012 |
| Intégrité de la recherche | 0,003 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,027 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».