MétaCan
Menu
Back to cohort
Record W3176883654 · doi:10.34778/1i

Correlational linkage analysis (Frequently Applied Designs)

2021· article· en· W3176883654 on OpenAlexaboutno aff
Laia Castro, Theresa Gessler, Sílvia Majó-Vázquez

Bibliographic record

VenueDOCA - Database of Variables for Content Analysis · 2021
Typearticle
Languageen
FieldSocial Sciences
TopicSocial Media and Politics
Canadian institutionsnot available
Fundersnot available
KeywordsLinkage (software)NewspaperFraming (construction)Context (archaeology)Content analysisData collectionSocial mediaPsychologyEconometricsComputer scienceData scienceSociologyGeographyMathematicsMedia studiesSocial scienceWorld Wide Web

Abstract

fetched live from OpenAlex

Correlational or second-order linkage analyses (Schulz, 2008) correlate content data points and survey data at the aggregate level. They are generally used to infer the impact of public opinion climate, the media context or media use on individual attitudes, cognitions and behaviors. Correlational linkage analyses make use of data collected at different points in time to be able to describe patterns of change and stability over time and to compensate for the reduced number of observations resulting from aggregating individual-level data. They often employ manual and automated content analysis, descriptive and inferential statistical analyses, and time series analysis.
 
 Field of application/theoretical foundation:
 Linkage analyses have extensively been used in the fields of political communication (Soroka, 2002), EU studies (Brosius et al., 2019a), and more recently, social media and social movements. Studies that employed second-order linkage analyses are related to theories of agenda setting (McCombs & Shaw, 1972), framing (Vliegenthart et al., 2008), or media bias and tone (Brosius et al., 2019b) (see chapter Content Analysis in Mixed Method approaches for a detailed account of applications and advantages of using linkage analyses).
 
 Example studies:
 In this data entry we describe two studies that regress survey data on media content data with additional weighs to better model news media effects. The first study (Boomgaarden & Vliegenthart, 2007) weigh media coverage of a particular topic (immigration) by issue prominence and circulation of the newspapers considered in the study. The second one (Vliegenthart et al., 2008) further introduces a publication recency moderator to account for how close in time a given news story was published from when survey data was collected and individuals may have been exposed to such piece of information.
 
 References
 Boomgaarden, H. G., & Vliegenthart, R. (2007). Explaining the rise of anti-immigrant parties: The role of news media content. Electoral Studies, 26(2), 404–417. https://doi.org/10.1016/j.electstud.2006.10.018
 Brosius, A., van Elsas, E. J., & de Vreese, C. H. (2019a). Trust in the European Union: Effects of the information environment. European Journal of Communication, 34(1), 57–73.
 Brosius, A., van Elsas, E. J., & de Vreese, C. H. (2019b). How media shape political trust: News coverage of immigration and its effects on trust in the European Union. European Union Politics, 20(3), 447–467. https://doi.org/10.1177/1465116519841706
 McCombs, M. E., & Shaw, D. L. (1972). The agenda-setting function of mass media. Public Opinion Quarterly, 36(2), 176–187.
 Schulz, W. (2008). Content analyses and public opinion research. The SAGE Handbook of Public Opinion Research, 348–357.
 Soroka, S. N. (2002). Issue attributes and agenda-setting by media, the public, and policymakers in Canada. International Journal of Public Opinion Research, 14(3), 264–285.
 Vliegenthart, R., Schuck, A. R., Boomgaarden, H. G., & De Vreese, C. H. (2008). News coverage and support for European integration, 1990–2006. International Journal of Public Opinion Research, 20(4), 415–439.
 
 Table 1. Data matching in correlational linkage analyses
 
 
 
 
 
 Author(s)
 
 
 Relationship of theoretical interest
 
 
 Sample
 
 
 Time frame
 
 
 Content-analytical constructs
 
 
 Linkage strategy
 
 
 
 
 Boomgarden & Vliegenthart (2007)
 
 
 News media reporting about immigration-related topics on aggregate share of vote intention for anti-immigrant parties
 
 
 (a) 157,968 articles collected through computer-assisted analysis, dealing with immigration and
 published in the five most-read Dutch national newspapers
 
 (b) Monthly self-reports on vote intention toward anti-immigrant parties from surveyed representative samples of the Dutch population
 
 (c) Monthly number of
 people that moved to the Netherlands and unemployment
 rates available from the Dutch governmental statistical institute
 
 
 1990-2002
 
 
 Visibility of immigration-related topics in news
 
 
 (1) The authors calculate a visibility score per article by computing:
 
 (1.1.) an average person’s log probability that
 s/he is exposed to news about immigration through a given article. This is done by using the frequency with which this article mentions immigration-related topics (f(t,a), both in the headline (fh(t,a)), in which case the frequency is weighed by 8, and in the body of the text (fb(t,a)), in which case the frequency is multiplied by 2. 
 
 (1.2.) 1.1. is weighed by circulation of the newspaper where the article is published (c(a)).
 
 (1.3.) 1.1. is weighed by whether the article is placed in the front page or other to account for how prominently the topic is featured (fp(a)).
 
 Notationally, the equation can be written as follows:
 (…)
 (2) In a second step, V(a) are aggregated for all articles in all outlets by month (the time unit to link content and survey data)
 
 (3) Final immigration visibility scores (independent variable) are linked to monthly percentage of people that reported intending to vote for an anti-immigration party (dependent variable) through time series analysis. The authors run ARIMA models, successively adding controls for extreme right leadership peaks (Fortuyn’s entrance in the political arena and assassination), immigration levels, unemployment rates, the interaction between the both and finally, the media visibility variables.
 
 
 
 
 Vliegenthart, Schuck, Boomgaarden, De Vreese (2008)
 
 
 How framing of EU news in terms of benefit and conflict explains public support for the EU
 
 
 (a) 329,746 articles that contained at least one reference to the European institutions in main newspapers of 7 EU countries (Denmark, Germany, Ireland, Italy, the Netherlands, Spain, and the United
 Kingdom) were computer-assisted content analysed to obtain data on EU media visibility.
 
 (b) 9,649 hand-coded articles that mentioned the EU at least twice (at least one of these references in
 the headline or in the lead of the article) were then analysed to investigate the framing of the EU. Approximately 50 articles per country were coded for each 6-month period.
 
 (c) Self-reports on EU support from the bi-annual standard Eurobarometer.
 
 
 1990–2006
 
 
 (a) News media attention/visibility of the EU
 (b) Presence of a benefit frame or a disadavantage frame in EU news coverage
 © Presence of a conflict framing in EU news coverage
 
 
 (1) Articles dealing with the EU (at least one reference) are weighed by prominence and publication recency as follows: Articles on the first page of a newspaper are counted twice as heavily as articles in the remainder of the newspaper; articles appearing in the month before a Eurobarometer survey was conducted are weighed six times, they are counted five times if appeared 2 months before, etc. The weighted EU visibility score is aggregated for each time period t in each country c.
 
 (2) Framing scores are then assigned to each article (benefit and disadvantage frames 0-2, conflict framing ranged from 0 to 3)
 
 (3) Mean framing scores per time period–country combination (fs(t,c)) are multiplied by visibility scores (vs(t,c)) to capture the overall salience of the frames (beyond its presence) as follows:
 (…)
 
 (4) OLS regressions with panel corrected standard errors are run with benefit, disadvantage and conflict framing as main independent variables, and aggregated-level support for the EU as dependent variable
 
 
 
 
 

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.851
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.004
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.116
GPT teacher head0.341
Teacher spread0.224 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueDOCA - Database of Variables for Content AnalysisSame topicSocial Media and PoliticsFrench-language works237,207