Unfaithful findings: identifying careless responding in addictions research
Bibliographic record
Abstract
Online data collection is inherently prone to careless responding and error. The growing use of web-based surveys within addictions research requires a greater understanding of how careless responding is defined and detected. Researchers are urged to place a greater emphasis on inspecting data and reporting data cleaning techniques. The quality of data gathered from self-reported measures is heavily dependent upon respondents' comprehension and motivation to self-disclose accurately. Although research methodologies strive to minimize the likelihood of misrepresentation, data collection methods that facilitate anonymity (e.g. computerized or web-based surveys) can increase the probability of careless responding 1, 2. Contrary to the common assumption that careless responding seldom happens and is unlikely to threaten data integrity, its prevalence has been reported to be as high as 40% 3, and rates of merely 5% have been shown to exaggerate or mute associations found between variables 2, 4. Such distorted effect sizes can increase the probability of Type I or II errors, especially among outcome measures with inherently high or low base rates 2, 5. This phenomenon, observed most extensively in personality research, has prompted the development of empirical methods for detecting careless and random responding. Within addictions research, however, its prevalence and impact upon research outcomes has remained largely unexplored. With the growing use of computerized and web-based surveys to assess addictive behaviors 6, a discussion of what careless responding is and how it can be detected within addictions research is necessary. Distinct from faking (i.e. responding deceptively), careless responding is characterized by participants' effortless or inattentive response behavior. Originally coined random responding, it is the tendency to respond to items without attention to content. It is generally assumed that such responses are truly random (i.e. equally likely to be chosen) and can be treated empirically as such 7; however, numerous scholars have argued that unmotivated response styles can be patterned and/or consistent. Consequently, many have chosen to use more descriptive terminology such as careless responding or insufficient effort responding, whereas others consider non-random responses as an extension or a subgroup of random responding, termed effective random responding. Definitions also vary, with some focusing on the lack of motivation or attentiveness in providing responses, while others define it as general psychological disengagement that may or may not be purposeful 2, 8-11. Careless responding has also been conceptualized as a subset of a much larger concept known as invalid responding 1. As definitions inform the development of techniques to identify problematic responses, such inconsistencies across the literature may explain why various rates are reported across studies. Despite the discrepancies, most definitions concur that careless responders introduce error to data, and recommend that these participants be identified and removed. Indeed, statistical handbooks recommend screening data visually for errors (e.g. out-of-range values, suspicious patterns) prior to data analysis 12, 13; however, more empirical and systematic detection methods exist. These can be organized into two types: (i) within-measure strategies that embed detection items/scales into tools and (ii) post-hoc strategies that employ statistical procedures to detect patterns and inconsistencies within data. The most popular within-measure techniques embed bogus items (i.e. obvious or nonsensical) into surveys and assume incorrect responses indicate inattentiveness. Although this method is effective, some have argued that incorporating absurd items may trivialize participation 12. Similarly, techniques that use content-specific items (i.e. validity indices) have been criticized for inadequately detecting or overestimating careless responses 14, 15. While the utility of within-measure strategies for detecting careless responding within addictions research is acknowledged 11, such techniques may affect participation, are costly and require validation prior to their use 1. Alternatively, post-hoc strategies provide researchers with suitable non-invasive options, and these are the main focus hereafter. Numerous post-hoc detection techniques with various degrees of validity across different populations have been reviewed 1, 8, 16. Overall, we organized approaches into four categories: response time, response pattern, multiple outlier analyses and internal consistency. The response time technique assumes that careless responders complete survey items significantly faster than those who are motivated to respond accurately. Although the experimenter is responsible for determining appropriate cut-off times for identifying careless responding in individual items, sections and entire surveys, a minimum of 2 seconds per item has been recommended previously 8. Alternatively, response pattern techniques assume that inattentive participants select the same response option(s) repeatedly. A well-known technique termed LongString 17 suggests that careless responding within Likert scales can be identified by establishing cut-points for acceptable response recurrence carefully; cut-points can be estimated by examining response option frequency curves or normative data on repeat responses 8. Similarly, multiple outlier analyses expect careless responders to deviate consistently from the sample norm. The Mahalanobis D statistic is one method of calculating the multivariate distance (i.e. D) between a respondent's scores and the sample mean across multiple items. Higher D values indicate a greater overall deviation from the sample, and as the square value of this index (i.e. D2) follows a χ2 distribution, empirical cut-offs (e.g. P < 0.001) can be used to identify careless responders 9, 18. In contrast, internal consistency strategies presume that unmotivated participants display great internal variability across survey data. Two of the most notable internal consistency methods include Goldberg's psychometric antonyms/synonyms 1 and Jackson's individual reliability index (IRI) 19. While both techniques compute an index score by correlating tool items, Goldberg's approach computes this score using empirically matched items–pairs and Jackson's IRI is calculated by correlating the split-halves of a tool (e.g. odd versus even items). Index scores closer to zero indicate careless responding for both techniques 1, 8, 9. Despite the seemingly vast availability of techniques for detecting careless responding, discretion should be exercised by researchers when selecting data cleaning strategies 10. Various factors, including questionnaire length, response type (e.g. scale, nominal), missing data and sample size, can limit which techniques are appropriate. None the less, their utility in improving the accuracy of data is undisputable, especially for online research where anonymity and possible compensation can invite participants to respond carelessly, quickly and dishonestly. With online surveys playing an increasingly larger role in addictions research, particularly among hidden substance-using populations 6, researchers are encouraged to become acquainted with techniques for detecting careless responding, their limitations and consequences of use. The confidence researchers and readers have in study findings can be improved further only through rigorous data inspection and transparent reporting. None.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.002 | 0.005 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".