MétaCan
Menu
Back to cohort
Record W4401506443 · doi:10.1073/pnas.2308950121

GPT is an effective tool for multilingual psychological text analysis

2024· article· en· W4401506443 on OpenAlexfundno aff
Steve Rathje, Dan-Mircea Mirea, Ilia Sucholutsky, Raja Marjieh, Claire Robertson, Jay Joseph Van Bavel

Bibliographic record

VenueProceedings of the National Academy of Sciences · 2024
Typearticle
Languageen
FieldSocial Sciences
TopicMisinformation and Its Impacts
Canadian institutionsnot available
FundersNational Institute of Mental HealthNatural Sciences and Engineering Research Council of CanadaTempleton World Charity FoundationSage FoundationRussell Sage Foundation
KeywordsComputer scienceNatural language processingArtificial intelligenceCoding (social sciences)Text miningSentiment analysisInformation retrieval

Abstract

fetched live from OpenAlex

The social and behavioral sciences have been increasingly using automated text analysis to measure psychological constructs in text. We explore whether GPT, the large-language model (LLM) underlying the AI chatbot ChatGPT, can be used as a tool for automated psychological text analysis in several languages. Across 15 datasets ( n = 47,925 manually annotated tweets and news headlines), we tested whether different versions of GPT (3.5 Turbo, 4, and 4 Turbo) can accurately detect psychological constructs (sentiment, discrete emotions, offensiveness, and moral foundations) across 12 languages. We found that GPT ( r = 0.59 to 0.77) performed much better than English-language dictionary analysis ( r = 0.20 to 0.30) at detecting psychological constructs as judged by manual annotators. GPT performed nearly as well as, and sometimes better than, several top-performing fine-tuned machine learning models. Moreover, GPT’s performance improved across successive versions of the model, particularly for lesser-spoken languages, and became less expensive. Overall, GPT may be superior to many existing methods of automated text analysis, since it achieves relatively high accuracy across many languages, requires no training data, and is easy to use with simple prompts (e.g., “is this text negative?”) and little coding experience. We provide sample code and a video tutorial for analyzing text with the GPT application programming interface. We argue that GPT and other LLMs help democratize automated text analysis by making advanced natural language processing capabilities more accessible, and may help facilitate more cross-linguistic research with understudied languages.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.026
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.026
Threshold uncertainty score0.086

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.026
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.002
Science and technology studies0.0010.000
Scholarly communication0.0020.004
Open science0.0020.004
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0260.019

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.084
GPT teacher head0.445
Teacher spread0.362 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations235
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueProceedings of the National Academy of SciencesSame topicMisinformation and Its ImpactsFrench-language works237,207