Whose Crystal Ball to Choose? Individual Difference in the Generalizability of Concept Testing
Bibliographic record
Abstract
The product development literature has identified several individual characteristics that could influence how subjects respond to new products in concept tests. Few of these characteristics have been thoroughly investigated. The purpose of this research is to examine whether a number of personality traits (1) do influence concept evaluation scores and (2) can be used to identify respondents who provide substantially higher‐quality data in concept testing and whether the answers to these questions change for major versus minor innovations. The data quality of the concept testing data is defined using the generalizability theory, which provides a decision‐specific G‐coefficient. Higher quality means a G‐coefficient closer to 1 for a particular managerial decision. A Web‐based study to concept test 10 appliance innovations on multiple occasions was conducted among 105 panelists from the Institute for Online Consumer Studies (IOCS). During the concept testing, respondents' innovativeness, change‐seeking tendency, and propensity to exert cognitive effort were also measured. The results showed that the respondent characteristics influence the mean evaluation of the concepts and the psychometric quality of the concept testing data: (1) there is a significant linear relationship between concept scores and all of the innovativeness scales and change‐seeking measures; (2) the effect of innovativeness on concept testing outcomes is even more substantial for major innovations than for minor innovations; (3) the study provides evidence that the quality of concept testing data provided by respondents varies substantially with their innovativeness, whereas the differences are more modest when scaling just minor innovations; (5) there are also strong effects on data quality for the Need to Evaluate scale used to capture cognitive effort characteristics; and (6) there is little effect of segmenting on social desirability on data quality. Managerially, the current results indicate that a product manager wanting to concept test a pool of appliance concepts can benefit from screening for the respondents who will provide higher‐quality concept testing data. For example, respondents who are high on domain‐specific innovativeness provide the highest‐quality concept testing data for both minor and major innovations. The effects of traits are stronger for major innovations, supporting the claim that subject selection is a more critical issue in concept testing of major innovations. Product managers can improve the quality of their concept testing data without an increase in cost by screening the subjects they use in concept testing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".