Proceedings of the Second Workshop on Computational Modeling of People’s Opinions, Personality, and Emotions in Social Media
Bibliographic record
Abstract
The idea of organizing PEOPLES stemmed from two related observations, namely the availability of large amounts of spontaneous data covering a range of personal aspects and the fact that such aspects are usually studied in isolation.Social media users nowadays freely express what is on their mind at any moment in time, at any location, and about virtually anything.These large amounts of spontaneously produced texts open up a unique opportunity to learn more about such users, e.g., predicting demographic variables (age, gender), but also personality types, as well as emotions and opinion expressions.This observation is not new, of course, and this opportunity has largely been exploited in the recent years, with abundant works on sentiment analysis, emotion detection, and personality.However, such traits of human personality and behavior have indeed attracted a substantial amount of attention but have been mostly studied in isolation, often in different -but related -communities, such as NLP, CL, AI.Therefore, we thought that the time was ripe to bring these communities a step closer to study people's traits and expressions jointly and in their interplay on such large volumes of available data.The communities' response, with 25 received submissions coming from 11 different countries and going well beyond typical NLP topics, proves again this year that there is wide interest at this intersection, and we are happy to be able to provide a context for exchanging ideas.Following the reviewers's advice, 14 papers were selected for inclusion in the proceedings.They cover a wide range of topics related to the three main PEOPLES themes (personality, emotion and opinion), their interaction and the impact of their modeling on social aspects like well-being, political preferences, humor and language use.To further enrich this volume, we additionally invited our keynote speakers to submit position papers that accompany their talks, and are excited that both of our keynotes submitted excellent papers touching upon issues of making NLP models more demographically aware and how researchers from related fields such as demography can benefit from NLP techniques.We hope that this is just the second edition of what will become series of workshops bringing together researchers in Computational Linguistics, Natural Language Processing and Computational Social Science, who share an interest in personality, opinion and emotion detection, and especially in researching the intertwining of such traits and expressions.We would like to thank our program committee consisting of 33 researchers from a variety of backgrounds for their insightful and constructive reviews.Without their support, this workshop would not have been possible.In addition, we thank all authors for submitting papers and making PEOPLES a big success.Also thanks to our two invited speakers, Dirk Hovy and Letizia Mencarini (Bocconi University, Italy), for having accepted to come to the workshop and share their expertise and ideas on PEOPLES' topics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.015 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".