MétaCan
Menu
Back to cohort
Record W4394978470 · doi:10.1101/2024.04.18.24306035

The Validity and Reliability of Dichotomized Self-rated Health Under Different Cutpoints

2024· preprint· en· W4394978470 on OpenAlexafffundabout
Charles Plante, Sharalynn Missiuna, Cory Neudorf

Bibliographic record

VenuemedRxiv · 2024
Typepreprint
Languageen
FieldSocial Sciences
TopicHealth disparities and outcomes
Canadian institutionsUniversity of SaskatchewanSaskatchewan Health Authority
FundersCanadian Institutes of Health Research
KeywordsLikert scaleScale (ratio)Reliability (semiconductor)PsychologyPublic healthHealth careCommunity healthApplied psychologyMedical educationMedicineNursingGeographyPolitical scienceCartography

Abstract

fetched live from OpenAlex

Abstract Self-rated health is a widely used indicator of overall health status. It is most often reported on a Likert scale of three to five values in surveys. To facilitate presentation and interpretation, it is common practice to simplify the variable by dichotomizing it; however, there has been little documented reflection on how this should be done. This paper explores all four possible dichotomizations of self-reported health, taken from three years of the Canadian Community Health Survey and reported by a Likert scale. We evaluated each dichotomization stratified by sociodemographic variables. We use regression analysis to explore the validity and reliability of all four possible dichotomizations by mapping them to the Health Utility Index. We found that lower cutpoints of dichotomization capture more pronounced differences in health status and are more consistent across sociodemographic variables. However, higher cutpoints of dichotomization should be considered for small data sets. About the Research Department The Saskatchewan Health Authority Research Department leads collaborative research to enhance Saskatchewan’s health and healthcare. We provide diverse research services to SHA staff, clinicians, and team members, including surveys, study design, database development, statistical analysis, and assistance with research funding. We also spearhead our own research programs to strengthen research and analytic capability and learning within Saskatchewan’s health system. About the UPHN The Urban Public Health Network (UPHN) is a national organization established in 2004 which today includes the Medical Officers of Health in 24 of Canada’s large urban centres. Working collaboratively and with a collective voice, the network addresses public health issues that are common to urban populations. Research operations of the UPHN are conducted in partnership with the University of Saskatchewan. Disclaimer This working paper is for discussion and comment purposes. It has not been peer-reviewed nor been subject to review by Research Department staff or executives. Any opinions expressed in this paper are those of the author(s) and not those of the Saskatchewan Health Authority. Suggested Citation Charles Plante, Sharalynn Missiuna, and Cordell Neudorf. 2024. “The Validity and Reliability of Dichotomized Self-rated Health Under Different Cutpoints.” medRxiv. Extended Abstract Introduction Self-rated health is a widely used indicator of overall health status. It is most often reported on a Likert scale of three to five values in surveys. To facilitate presentation and interpretation, it is common practice to simplify the variable by dichotomizing it; however, little documented reflection has been done on how this should be done. Methods We use regression analysis to explore the validity and reliability of all four possible dichotomizations of self-reported health in the Canadian Community Health Survey in 2013-2015 by mapping them to a validated health measure: the Health Utility Index Mark 3 (HUI). We posit that more valid cutpoints in self-rated health are associated with larger changes in HUI. We posit further that more reliable cutpoints are associated with similar changes across sociodemographic variables, including age, sex, education, marital status, geography and income. We also provide descriptive statistics to contextualize our analysis. Results The greatest proportion of respondents reported having “very good” health, although the proportion of the population reporting “excellent” or “very good” health decreased with age. Similarly, Canadians tend to score highly in HUI. Our regression results suggest that HUI tends to be higher for younger, richer, married, educated and urban populations. However, these associations are muted as the cutpoint used to dichotomize self-reported health is raised. The model with the lowest cutpoint, distinguishing between poor health and all other health statuses, was associated with the greatest and most consistent negative changes in HUI among different sociodemographic groups. Conclusions Dichotomizing self-rated health using lower cutpoints captures more pronounced differences in health status measured by HUI and tends to capture more consistent differences across sociodemographic variables. That is, lower cutpoints produce more valid and reliable results. However, lower cutpoints isolate less commonly reported health levels and may lead to less accurate results in smaller populations. Key Points This article addresses the knowledge gap concerning the most accurate way to dichotomize self-rated health data reported using a Likert scale. This paper explores the validity and reliability of all four possible dichotomizations of self-reported health reported by a Likert scale. Lower cutpoints of dichotomization capture more pronounced differences in health status and are more consistent across sociodemographic variables. Higher cutpoints of dichotomization should be considered for small data sets.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.078
Threshold uncertainty score0.994

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0050.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0000.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.049
GPT teacher head0.370
Teacher spread0.320 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations12
Published2024
Admission routes3
Has abstractyes

Explore more

Same venuemedRxivSame topicHealth disparities and outcomesFrench-language works237,207