MétaCan
Menu
Back to cohort
Record W2192884026 · doi:10.1111/add.13221

Unfaithful findings: identifying careless responding in addictions research

2015· editorial· en· W2192884026 on OpenAlexaff
Alexandra Godinho, Vladyslav Kushnir, John Cunningham

Bibliographic record

VenueAddiction · 2015
Typeeditorial
Languageen
FieldPsychology
TopicBehavioral Health and Interventions
Canadian institutionsCentre for Addiction and Mental Health
Fundersnot available
KeywordsPsychologyAddictionMisrepresentationData collectionAnonymityComprehensionData qualitySocial psychologyComputer scienceComputer securityStatisticsPsychiatry

Abstract

fetched live from OpenAlex

Online data collection is inherently prone to careless responding and error. The growing use of web-based surveys within addictions research requires a greater understanding of how careless responding is defined and detected. Researchers are urged to place a greater emphasis on inspecting data and reporting data cleaning techniques. The quality of data gathered from self-reported measures is heavily dependent upon respondents' comprehension and motivation to self-disclose accurately. Although research methodologies strive to minimize the likelihood of misrepresentation, data collection methods that facilitate anonymity (e.g. computerized or web-based surveys) can increase the probability of careless responding 1, 2. Contrary to the common assumption that careless responding seldom happens and is unlikely to threaten data integrity, its prevalence has been reported to be as high as 40% 3, and rates of merely 5% have been shown to exaggerate or mute associations found between variables 2, 4. Such distorted effect sizes can increase the probability of Type I or II errors, especially among outcome measures with inherently high or low base rates 2, 5. This phenomenon, observed most extensively in personality research, has prompted the development of empirical methods for detecting careless and random responding. Within addictions research, however, its prevalence and impact upon research outcomes has remained largely unexplored. With the growing use of computerized and web-based surveys to assess addictive behaviors 6, a discussion of what careless responding is and how it can be detected within addictions research is necessary. Distinct from faking (i.e. responding deceptively), careless responding is characterized by participants' effortless or inattentive response behavior. Originally coined random responding, it is the tendency to respond to items without attention to content. It is generally assumed that such responses are truly random (i.e. equally likely to be chosen) and can be treated empirically as such 7; however, numerous scholars have argued that unmotivated response styles can be patterned and/or consistent. Consequently, many have chosen to use more descriptive terminology such as careless responding or insufficient effort responding, whereas others consider non-random responses as an extension or a subgroup of random responding, termed effective random responding. Definitions also vary, with some focusing on the lack of motivation or attentiveness in providing responses, while others define it as general psychological disengagement that may or may not be purposeful 2, 8-11. Careless responding has also been conceptualized as a subset of a much larger concept known as invalid responding 1. As definitions inform the development of techniques to identify problematic responses, such inconsistencies across the literature may explain why various rates are reported across studies. Despite the discrepancies, most definitions concur that careless responders introduce error to data, and recommend that these participants be identified and removed. Indeed, statistical handbooks recommend screening data visually for errors (e.g. out-of-range values, suspicious patterns) prior to data analysis 12, 13; however, more empirical and systematic detection methods exist. These can be organized into two types: (i) within-measure strategies that embed detection items/scales into tools and (ii) post-hoc strategies that employ statistical procedures to detect patterns and inconsistencies within data. The most popular within-measure techniques embed bogus items (i.e. obvious or nonsensical) into surveys and assume incorrect responses indicate inattentiveness. Although this method is effective, some have argued that incorporating absurd items may trivialize participation 12. Similarly, techniques that use content-specific items (i.e. validity indices) have been criticized for inadequately detecting or overestimating careless responses 14, 15. While the utility of within-measure strategies for detecting careless responding within addictions research is acknowledged 11, such techniques may affect participation, are costly and require validation prior to their use 1. Alternatively, post-hoc strategies provide researchers with suitable non-invasive options, and these are the main focus hereafter. Numerous post-hoc detection techniques with various degrees of validity across different populations have been reviewed 1, 8, 16. Overall, we organized approaches into four categories: response time, response pattern, multiple outlier analyses and internal consistency. The response time technique assumes that careless responders complete survey items significantly faster than those who are motivated to respond accurately. Although the experimenter is responsible for determining appropriate cut-off times for identifying careless responding in individual items, sections and entire surveys, a minimum of 2 seconds per item has been recommended previously 8. Alternatively, response pattern techniques assume that inattentive participants select the same response option(s) repeatedly. A well-known technique termed LongString 17 suggests that careless responding within Likert scales can be identified by establishing cut-points for acceptable response recurrence carefully; cut-points can be estimated by examining response option frequency curves or normative data on repeat responses 8. Similarly, multiple outlier analyses expect careless responders to deviate consistently from the sample norm. The Mahalanobis D statistic is one method of calculating the multivariate distance (i.e. D) between a respondent's scores and the sample mean across multiple items. Higher D values indicate a greater overall deviation from the sample, and as the square value of this index (i.e. D2) follows a χ2 distribution, empirical cut-offs (e.g. P < 0.001) can be used to identify careless responders 9, 18. In contrast, internal consistency strategies presume that unmotivated participants display great internal variability across survey data. Two of the most notable internal consistency methods include Goldberg's psychometric antonyms/synonyms 1 and Jackson's individual reliability index (IRI) 19. While both techniques compute an index score by correlating tool items, Goldberg's approach computes this score using empirically matched items–pairs and Jackson's IRI is calculated by correlating the split-halves of a tool (e.g. odd versus even items). Index scores closer to zero indicate careless responding for both techniques 1, 8, 9. Despite the seemingly vast availability of techniques for detecting careless responding, discretion should be exercised by researchers when selecting data cleaning strategies 10. Various factors, including questionnaire length, response type (e.g. scale, nominal), missing data and sample size, can limit which techniques are appropriate. None the less, their utility in improving the accuracy of data is undisputable, especially for online research where anonymity and possible compensation can invite participants to respond carelessly, quickly and dishonestly. With online surveys playing an increasingly larger role in addictions research, particularly among hidden substance-using populations 6, researchers are encouraged to become acquainted with techniques for detecting careless responding, their limitations and consequences of use. The confidence researchers and readers have in study findings can be improved further only through rigorous data inspection and transparent reporting. None.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Research integrity, Insufficient payload (model declined to judge)
Consensus categoriesResearch integrity, Insufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Editorial · Consensus signal: Editorial
Teacher disagreement score0.110
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0030.002
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0020.005
Insufficient payload (model declined to judge)0.0030.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.178
GPT teacher head0.514
Teacher spread0.336 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEditorial

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations42
Published2015
Admission routes1
Has abstractyes

Explore more

Same venueAddictionSame topicBehavioral Health and InterventionsFrench-language works237,207