Consumer insights from text analysis
Notice bibliographique
Résumé
Language, whether spoken or written, is fundamental to the consumer experience. It is how consumers express their thoughts, articulate choices, negotiate with others, and receive information about products or services. And it is how marketers deliver persuasion attempts, make apologies, and build relationships with consumers. Language has also long been a powerful research tool. Scholars have used content analysis methods like ethnography in interviews, observational studies, and interpretation of language-based artifacts to advance our understanding of consumer culture (e.g., Arnould & Thompson, 2005; Stern, 1989). Research on the psychology of language has studied how people respond to a range of semantic, syntactic, and rhetorical aspects of verbal communication (Kronrod, 2022). But something new has emerged in the last decade or so—something that has helped foster a genuine explosion in the analysis of text in consumer research (Packard & Berger, 2023). First, language data has become more accessible. While consumers, companies, and other marketplace actors are constantly producing language, only recently has much of this content become digitized (or digitizable), making it far easier to collect and analyze. Every day, billions of consumers share attitudes and opinions online. Customer service calls, depth interviews, and Zoom meetings can be transcribed with the push of a button, and the shift from paper and pencil surveys to online data collection means open-ended participant responses are ready-made for automated text analysis. Similarly, massive online repositories of human conversations, product reviews, books, movie scripts, newspaper articles, and other cultural content provide easy ways to explore ideas in language. Second, new tools have changed how language can be analyzed. Previously, language data could only be coded manually. Researchers, or research assistants, would read or listen to language and score it on various dimensions. While manual coding is helpful, it is often subjective and difficult to scale for both lab and field research. Manually reading and carefully evaluating just 10 conversations, online reviews, or thought listings takes a fair amount of time—and reading 1000 takes 100 times as long. In recent years, though, psychologists and computer scientists have developed tools that allow language data to be processed and analyzed quickly and easily. Dictionaries like Linguistic Inquiry and Word Count (LIWC, Tausczik & Pennebaker, 2010) allow researchers to count the presence of different words linked to psychological constructs and approaches like latent Dirichlet allocation (Blei et al., 2003). Word embeddings (cf. Bakarov, 2018) and large language model approaches (e.g., BERT, GPT) make it possible to measure almost any construct. And these tools are becoming more user-friendly every day. Just like the microscope revolutionized chemistry and the telescope revolutionized astronomy, ongoing developments in automated textual analysis have allowed researchers in a variety of different domains to unlock a range of new insights from text. This special issue presents a collection of exciting articles that showcase how text analysis can illuminate a diverse range of theoretical and substantive topics in consumer psychology. Through the set of theories explored, methods applied, and tools introduced, we hope that scholars who have just begun to explore the space, those who fear they do not have the “know how,” and even text analysis and psychology of language experts will be inspired to consider how these methods can generate new consumer insights. Next, we offer a brief synopsis of why each of the special issue articles can help do just that. We discuss articles that combine text analysis and experiments, introduce user-friendly ways to capture constructs and apply cutting-edge methods to novel problems. We then share a brief perspective on the special issue's methodological diversity, discuss who is using text analysis today, and point toward some text analysis resources. Text analysis does not have to be primary in a given investigation. It can complement a more traditional experimental or even qualitative effort. Two special issue articles offer examples of research that puts conceptual or substantive issues first, building on theory from regulatory mode, construal, and authenticity to explore important consumer phenomena. Pham, Septianto, Mathmann, Jin, and Higgins use a mixture of traditional experiments and automated text analysis of field data to demonstrate how construal and regulatory modes interactively influence social media sharing. Specifically, when consumers are presented with more abstract language, they can better integrate thought and action, which leads to heightened engagement with and re-transmission of the social media content. This article provides a novel demonstration of how text analysis can be used to shed new insights into some of the most broadly investigated theory spaces in psychology. Nam, Balakrishnan, De Freitas, and Brooks examine how consumers react when brands respond to sociopolitical events. They text analyze the sentiment of consumer comments on brand Instagram posts about the Black Lives Matter movement, finding that when brands post more quickly after specific events, consumer comments are more positive. The authors supplement this secondary data with experiments that explore when and why the speed of brand responses to sociopolitical events matters. They find that, regardless of what brands say, the speed with which they respond to events can influence the brand's perceived authenticity and consumers' purchase intentions. Well-validated dictionaries and tools are essential to measuring psychological constructs expressed in language. The special issue presents three articles that make such contributions. From a traditional dictionary approach to capturing nostalgia, to more sophisticated models (with simple front-end tools) that measure theory of mind and psychological distance, these articles offer new methods and preliminary insights as well as ideas for future research in related areas. Chen, Wildschut, Layous, and Sedikides develop and validate a dictionary to capture the degree of nostalgia expressed in a text. The authors outline multiple possibilities for how to use this dictionary, from exploring nostalgia in new contexts to providing a quick manipulation check, and discuss some potential insights that the dictionary offers about nostalgia itself. Hartmann, Bergner, and Hildebrand demonstrate that attribution of mind to smart objects, as observed in consumer language, moderates consumer relationships with and usage of these devices. Consumers who attribute mind to smart objects express a more communal (vs. instrumental) relationship with the object, which leads them to use it in richer, more diverse ways. To explore such relationships, the authors present MindMiner, a user-friendly web-based tool that scholars can use with their own data to capture theory of mind in natural language. Sepehri, Mirshafiee, and Markowitz's PassivePy tool offers a sophisticated, accurate, and newly accessible method to help researchers capture verb voice. While verb voice (active vs. passive voice) decisions are made for every sentence and have important psychological and marketing-relevant correlates (e.g., psychological distance, attribution), it has proven very difficult to capture this complex grammatical feature automatically. The article introduces this new tool while contributing novel insights on how consumer use of passive voice might be linked to consumer complaining, online reviews, and charitable giving. The special issue also showcases more advanced natural language processing methods and models. From the increasingly common use of topic modeling to more custom word embeddings and neutral networks, these articles show how capturing elusive constructs like bias, subjectivity, and style can be achieved by analyzing language collected in the field or experimentally. Rathee, Banker, Mishra, and Mishra use word embeddings and field experiments to show that algorithms learn gender bias from language. In particular, algorithms associate females with more negative traits than males (e.g., irrational, impulsive). Critically, the authors demonstrate that this bias fundamentally changes females' consumption experiences. Relative to males, females are more likely to be served with negatively biased ads and search engine results. Ultimately, this alters—and constrains—their consideration sets and product choices. Park, Song, and Sela use a convolutional neural network to examine how subjective and objective content contributes to the helpfulness of online reviews. While one might imagine that objective content is more helpful, it turns out the subjective information is helpful as well. That said, while each type of content is helpful on its own, combining the two actually reduces helpfulness slightly. Boghrati, Berger, and Packard combine old and new text analysis methods to test one of the oldest questions in communication—beyond the content of communication, how much does communication style matter? Using measures from one of the most widely used dictionaries (LIWC), along with more sophisticated methods from machine learning (support vector machines) and topic modeling, they isolate style from content and explore its unique contribution to impact for tens of thousands of academic articles. Experiments then help shed light on specific ways communicators can use style to shape an idea's success. Wan, Xu, and Kuchmaner use an expectation disconfirmation framework to examine the content of food reviews. Specifically, they demonstrate that while postpurchase product performance has an important impact on review writing, the prepurchase phase also plays a role. Further, expectations about a product's nutritional value and disconfirmation of these expectations shape review content. The paper's modeling approaches use a complex assortment of text analysis methods including latent Dirichlet allocation (LDA) and measures of lexical diversity, readability, sentiment, and subjectivity. Taken together, the articles in this special issue spotlight the rich diversity of ways in which language can be analyzed to shed light on consumer psychology and behavior. Language is used in the special issue as a predictor (four articles), an outcome measure (two articles), or is revealed to have potential applications as both (three articles). Text analysis methods covered include simple sentiment analysis and widely used LIWC applications (four articles), custom-developed dictionaries and tools (three articles), and highly sophisticated language models (four articles). Several articles use a combination of such methods (four articles). Figure 1 locates the special issue articles in terms of their sources of data (experimental vs. field) and an estimate of the relative ease of use of the text analysis approaches applied. We note that more than half of the articles incorporate data from experiments conducted in the lab or field, challenging the notion that text analysis is primarily suited to archival data. For example, Rathee et al. (2023) run clever field experiments using real online advertising campaigns to test their ideas about how ad algorithms learn and perpetuate negative gender biases. Further, we observe no clear correlation between the use of field data and more sophisticated natural language processing (NLP) methods among the special issue articles: using archival field data does not necessarily require advanced methods. Special issue articles used methods that ranged from relatively straightforward applications of easy-to-use sentiment analysis and dictionaries (e.g., Nam et al., 2023; Wan et al., 2023) to more complex methods like grammar dependency parsing or deep learning neural network approaches (e.g., Park et al., 2023). Such methods are applied when doing so help solve specific research challenges, such as the difficulty of accurately measuring passive voice (Sepehri et al., 2023) or subjectivity in language (Park et al., 2023). It is worth noting that several of the special issue authors would not necessarily be thought of as “text analysis people.” Some may have recently added the method to their toolkit. Others partnered with colleagues that brought know-how to the team. This raises an interesting question: who's using text analysis and how? To find out, we asked people on the Society for Consumer Psychology and Association for Consumer Research email lists (research professors and PhD students; N = 220) to answer some questions about their research methods. We did not mention text analysis to try to mitigate self-selection. Results indicate that over two-thirds (69%) of scholars—both junior and senior—reported using automated text analysis in their research at least once, suggesting that it is indeed approaching mainstream status as a research method. Most researchers who noted they had not used automated text analysis yet said it was primarily due to a lack of knowledge (80%). Hopefully, this special issue will help more scholars see the opportunity and ease of adding automated text analysis to their tool kit or to their research team by finding a colleague to join the effort. The articles in the special issue and the use cases outlined in our survey highlight several ways consumer researchers can use text analysis to explore almost any topic. First, researchers can use text analysis tools to offer richer tests of existing ideas. They may already have a particular idea in mind about a relationship between a predictor and outcome, or they may even have a sense of what the underlying process might be. Further, they may have conducted some experiments to test these ideas. Automated text analysis can provide further tests, uncover additional relationships, or provide external validity. In experimental work, for example, researchers can have participants respond to open-ended questions or write about experiences, and parse that content to assess dependent variables (Barasch & Berger, 2014; Spiller & Belogolova, 2017), mediators (Wu et al., 2019), or alternative explanations. Similarly, researchers who have conducted carefully controlled experiments may want to test whether their effects hold in the noisy field. Online reviews (Lafreniere, Moore, & Fisher, 2022), social media posts (Lee & Junqué de Fortuny, 2022), and everything from newspaper articles and books to movie scripts and song lyrics (Packard & Berger, 2020; Toubia et al., 2021) can provide useful testing grounds for external validity. Second, researchers can use these approaches to help develop ideas and theories in the first place (van Osselaer & Janiszewski, 2021). Some researchers might be interested in a broad question, like what makes online content viral, or what about customer service interactions lead to greater customer satisfaction. In such situations, they may have a dependent variable in mind but are not sure which independent variable to focus on. By using automated textual analysis to measure multiple independent variables, researchers can explore which features matter most (Hodges et al., 2023), and use that to decide where to focus before developing theory and designing subsequent experiments. Similarly, researchers can test alternative explanations by measuring and controlling for other factors that may play a role; they can use automated textual analysis to explore potential underlying processes, simultaneously testing them. Third, researchers can use the findings of text analysis to improve their academic writing. Recent work has used natural language processing to explore how to make writing clearer and why some articles are cited more (Warren et al., 2021); Boghrati et al., 2023). Abstraction, technical language, and passive writing can make research difficult to understand and lead it to be cited less. Similarly, writing more simply, using present (rather than past) tense, and using personal voice (e.g., “we” rather than “results” find) can all be useful to improve impact. Finally, we note that this special issue offers only a sample of the great research out there that analyzes text for consumer insight. Exploring other text analysis articles might help scholars further understand the diversity of approaches and research problems that these tools might help solve. Recent review articles offer many examples (Kronrod, 2022; Packard & Berger, 2023). For researchers interested in trying out these approaches, there are simple ways to begin exploring. Start by getting some text; this can be responses to a writing prompt from an experiment, thought listings, or online reviews. Then put this data in a spreadsheet, where each participant or observation is a row. Next, try inputting that data into some of the free online tools available for extracting features, such as http://textanalyzer.org/ or http://www.lexicalsuite.com/. LIWC (https://www.liwc.app/) also allows one to try inputting text before purchasing the software. For researchers who want to try more sophisticated approaches, a variety of recent papers provide helpful direction (Berger & Packard, 2022; Boyd et al., 2021; Humphreys & Wang, 2018). Finding experienced coauthors and free online courses or websites can offer details on specifics and more advanced methods. Finally, once you have got some basics on how to do automated text analysis, reading up on the psychology of language itself could offer ideas about what to measure, and how your findings might fit into language research more broadly. In addition to the review papers cited earlier, books from various traditions provide a great overview of the mechanics and characteristics of language (e.g., Berger, 2023; Fahnestock, 2011; Pinker, 2007). The analysis of unstructured data, using automated tools, promises to revolutionize the social sciences. Our hope is that this special issue's collection of novel demonstrations of using text analysis for insights helps inspire more consumer psychologists to help lead this revolution. From understanding how advertising contributes to gender bias (Rathee et al., 2023) to how theory of mind shapes smart object usage (Hartmann et al., 2023), and offering new ways to capture and understand core constructs such as construal, attribution, credibility, and psychological distance (Pham et al., 2023; Sepehri et al., 2023) shows that text analysis is no longer just about online reviews and social media posts. By understanding automated text analysis and the different ways it can be used, consumer psychologists can shed light on a range of interesting conceptual and substantive questions. The authors thank Lauren Block, Jen Argo, and Thomas Kramer for the opportunity to serve as editors of this special issue, and Lauren Block and Sandy Osaki for their sage guidance and support throughout the process. We also wish to thank the special issue reviewers, who were incredibly generous with their time and constructive comments across the surprisingly large number of submissions received. No data was collected.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,002 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».