Potential High-Dimensionality Structure in the 449-tda (Self and Peer Ratings)
Notice bibliographique
Résumé
ANALYSIS PLAN: Potential High-Dimensionality Structure in the 449-tda (Self and Peer Ratings) OSF Project Name: Further Analysis of 449-TDA Data BROAD RESEARCH QUESTION: How much does this lexical-personality-study data (the 449 trait-descriptive adjectives, which although not the largest is the most frequently cited source supporting a six-factor structure for English-language personality-trait adjectives) yield evidence supporting also ‘high-dimensionality’ structures of 15 and 20 factors? The high-dimensionality structures mentioned were derived from a more comprehensive selection of 1710 adjectives (although in a smaller set of self-ratings) and replicated in a reduced set of 540 adjectives (with both peer ratings and a larger set of self-ratings). The same question can be asked in two more specific forms. Question 1: When one extracts 15 or 20 factors in these data, how many of the expected factors appear? Question 2: When one determines the most robust ‘emic maximum’ factor structure for this data, how much do the factors in that structure overlap with the high-dimensionality structures of 15 and 20 factors? ‘Emic maximum’ refers to the structure having the largest number of interpretable factors that are sufficiently sized. ‘Robustness’ is here operationally defined as a tendency to appear regardless of data handling (ipsatized versus original data), use of orthogonal versus oblique rotation method, and use of self versus peer rating subsets of the data. The variable selections are all a priori rather than determined by the current investigators, and derived from frequency-of-use judgments in a Canadian setting, rather than based on work by Norman and Goldberg from 40-50 years ago, as in the previously studied sets of adjectives. 1. A PRIORI–SPECIFIED PROCEDURE FOR EXPLORATORY FACTOR ANALYSES 1. For the 449-tda (trait-descriptive adjective) set, with ipsatized data, initially combining self- and peer-rating data, apply two methods that recommend a specific number of factors (parallel analysis and the MAP-index based on PCA [based e.g. on Zwick and Velicer, 1986) to determine an initial range of ‘numbers of factors’ to consider – from the lowest recommended by any method to the highest recommended by any method/variable-set combination, as well as the best starting point (that number most redundantly recommended across these combinations, if any are redundant). Specifically, parallel analysis (PA) will be executed first, as it does not require initially setting a maximum number of factors; if PA recommends more than 33 factors, that number recommended by PA will be set of maximum number of factors, otherwise the default of 33 factors will be the initial set maximum for the MAP procedure (based on a previous study with the more comprehensive 1710-tda in which 33 factors was the maximum number of factors that was ever indicated by PA or MAP methods), although if no clear minimum has been obtained at 33 factors or less, a higher minimum will be sought. 2. Within that range, run varimax, equamax, and oblimin rotations (as in Goldberg, 1990) for various numbers of factors. Previous work with the 1710-tda indicated the latter two methods are best for identifying high-dimensionality structures, but varimax is also included because it has been the typical rotation method in studies of personality descriptors from the natural language. One prominent method of oblique rotation – promax – is intentionally omitted from the analyses because it is merely a variant of varimax and would tend to converge extremely highly (be redundant) with varimax. 3. Within the determined range, discard solutions that have any insufficiently sized factors, by a preset ‘sufficient size’ minimum of 3 salient terms with a loading of at least .3 in absolute magnitude, also including at least one salient term with a loading of at least .4 in absolute magnitude. These threshold numbers of salient terms are proportional in comparison to those used in a study with the much larger 1710-tda, in which the threshold was 8 salients of at least .3, including 3 with at least .4. 4. For the solutions that remain, determine the maximum number of interpretable factors. That is, any solutions that have any factors that are judged by both of the two investigators as impossible to interpret substantively are eliminated. 5. For each of the remaining solutions – the maximum number-of-factors (that are sufficiently sized and interpretable) within each data set are set out as a candidate structure, and then compared across data sets, so as to sort out which number-of-factors has relatively better and worse convergence across methods (i.e., use of ipsatized versus original/raw data, of orthogonal vs. oblique rotation, and of self- versus peer-rating subsets of the data). Presumably, just one of the three emic-maximum candidate structures (i.e., the one from varimax, from equamax, and from oblimin) will stand out as the most robust and the best emic-maximum model. Steps 1 through 5 are all oriented toward identifying the optimal ‘emic maximum’ model from these datam and thus relevant to Question 2 above. Steps 6 and 7 pertain to the various, more a priori, ‘etic’ models also used in this study’s comparisons, and thus relevant to Question 1. 6. To the extent that the above procedure did not examine solutions at the one- to six-factor level, and at the 15- and 20-factor level, add these, with comparison on the same three robustness indices. The one-factor level will be indexed by the first unrotated principal component. Two- through six-factor levels will be indexed by varimax-rotated components, consistent with the previous literature. Consistent with the method associated with each in the previous study, the 15-factor level will be indexed by an equamax rotation, and the 20-factor level by an oblimin rotation (delta=0). 7. Replication of the 15- and 20-factor structures in the 449-tda data will be enabled by use of marker terms for each factor determined from previous study results. That study identified 8-item adjective scales for each factor, based mechanistically on the highest loadings on each pole of the factor. That subset of each set of 8 items that is represented in the 449-tda will be used as marker-indicators for each factor, provided at least 3 items are found in the 449-tda. If fewer than 3 are found, factor-loading tables from the previous study will be consulted to identify enough additional terms based on the next highest loading(s), and is found also in the 449-tda and can be used as a marker-indicator; in this manner, each factor in the 15- and 20-factor models should have at least three marker-indicators. The sets of marker-indicators can be employed to gauge replication in two ways: (a) what percentage of the marker-indicators is associated with a unique factor in the 449-tda data, and (b) how much the scored composites of the marker-indicators correlate with the various 449-tda factors. By these indices a high degree of replication would involve (a) high percentages for each 449-tda factor, and for each set of marker-indicators, and a high-percentage for each 449-tda factor associated with one and only one of the marker-indicator sets, and (b) high convergent and low divergent correlations of the marker-indicator sets with the 449-tda factors. In either case, high degree of replication would involve one-to-one matches between marker-indicator sets and individual factors. As a useful comparison to results of this ‘import markers’ approach, a reverse “export markers’ approach will also be run. 15- and 20-factor structures from the 449-tda data will be indexed by mechanistically derived 8-item scales, and these will be scored in the datasets from the previous study, and compared to 15- and 20-factor structures there. Notes: (a) Orthogonal versus oblique comparisons will be varimax vs. oblimin for one-to six-factor solutions and for the emic-maximum varimax model. For all other comparisons, these will be equamax vs. oblimin comparisons, on the basis that as the number of factors becomes large these two methods have been observed to be able to produce more interpretable factors of sufficient size. (b) The methods for examining robustness in step 5 must vary according to the comparison: The ipsatized-original comparison can utilize canonical correlation analysis, the orthogonal-oblique comparison will need to utilize comparison of factor scores (the next most optimal method), and the self-peer comparison (because these are same variables but different cases) will need to rely on Tucker coefficients of factor congruence. (c) In this study, self- and peer-rating data are combined to maximize simplicity and statistical power, but the degree of replication between self- and peer-data is taken into account with the robustness analyses. (d) In this study, one-to six-factor models are examined only for comparison with high-dimensionality models with respect to robustness. (e) For comparison, we will also examine the parallel analysis and MAP results for original (non-ipsatized) data; these are expected to indicate slightly more factors, but previous study indicates that factors based on original data tend to be somewhat less robust than those based on ipsatized data. ADDITIONAL COMPARISONS INVOLVING INTERNAL CONSISTENCY AND FACTOR INDEPENDENCE Factor scores (i.e., component scores from principal components analyses) will be retained for each factor model derived from the 449-tda, but the prime comparisons will involve aggregates of adjectives selected as best representatives of the various factors; this is because factors could be orthogonal even if the most salient indicators for them are not orthogonal upon aggregation. For each factor, up to eight adjectives will be selected by the ‘highest-loading items’ method. For each of the opposing poles of the factor, the four adjectives with the highest loading will be selected, so long as the loading in ques
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,862 | 0,372 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».