A Critical Review of Measures of Mentalization from Peter Fonagy’s Conceptual Framework
Bibliographic record
Abstract
Background: Mentalization is an expansive and multifaceted construct with important implications for psychological research as well as the etiology and treatment of psychological disorders. It is defined as the process through which an individual infers their own and others’ mental states are intentional and lead to meaningful actions (Bateman & Fonagy, 2004). The theoretical framework in which mentalization is situated has evolved over the years and is widely accepted however, the empirical work surrounding the measurement and operationalization of mentalization is not as well developed and merits further investigation. Methods: The authors examined the psychometric soundness and construct validity of the gold standard, interview-based measure of mentalizing as well as five self-report scales. To assess the convergent validity of these measures, Canadian university students (N = 247) completed three self-report measures of mentalization as well as one task-based tool, the Movie for the Assessment of Social Cognition (MASC; Dziobek et al., 2006). To investigate whether self-report measures predict performance on the MASC, twenty linear regressions were estimated. Exploratory factor analysis was conducted to identify the common latent factors underlying all five self-report measures at the subscale level. Results: Certain self-report measures were strongly linked and common content included items about emotion recognition and regulation, understanding one’s motivations for their behaviors and making accurate inferences about others’ thoughts. Other measures were weakly correlated and dissimilar in item content. All self-report measures were weakly correlated with the MASC. Most of the regression models were non-significant. Of the four models that emerged as significant and had significant direct effects, a classical suppression effect was observed, which merits replication. Exploratory factor analysis revealed a one factor solution fit the subscale level data well. Conclusions: This study provides preliminary evidence that there is some convergence among self-report measures. Researchers should reflect on their choice of instrument and should not use different tools interchangeably. The lack of convergence between task-based and self-report measures is disconcerting and warrants further research in this area. Last, whether the latent construct of mentalization is indeed unidimensional or has a more complex factor structure is yet to be determined.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.025 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".