Self-reported Legal Status in the California Health Interview Survey: An evaluation of data quality and application towards adolescent mental health
Bibliographic record
Abstract
Legal status is an important social determinant of health for immigrants and children of immigrant parents, which is typically not measured in public health surveys. The sensitivity of legal status and presumed response behavior to relevant questions are primary reasons why this topic goes unmeasured. Changes in immigration enforcement likely impact the sensitivity of the topic and may compromise data quality, however, this is also likely when legal status matters most for health outcomes. This dissertation evaluates the response behavior to questions of citizenship and immigration status in the California Health Interview Survey and applies these data to identify mental health risks for Latino adolescents with an unauthorized parent.The first study, When we ask, do they answer? Item-nonresponse to questions of citizenship and immigration status in the California Health Interview Survey, examined foreign born survey participants who did not answer questions of citizenship and immigration statusbetween 2001 and 2015. Nonresponse was low overall, however, increased over time and was largely attributable to respondents who were born in Mexico. The second study, When they answer, should we listen? Examining the quality of self-reported citizenship and immigration status, evaluated potential misreporting of legal status among Mexican-born participants between 2003 and 2015. This study utilized indirect estimation strategies which have been developed to produce profiles of the unauthorized population from surveys which do not ask legal status. Nearly a quarter of all Mexican-born participants reported that they were a non-citizen without agreen card, and these participants were demographically similar to external profiles of the unauthorized population. Predicted probabilities of unauthorized status produced by the indirect estimation procedure indicated that the threat of extensive misreporting was low and consistent over time. These results, paired with the findings of low nonresponse, indicate that participants were willing to answer questions of citizenship and immigration status and that these data are fit for use. The third paper, Severe Psychological Distress Among Latino Adolescents with an Unauthorized Parent examined adolescent mental health using data from 2007 to 2016 disaggregated by parental nativity and legal status. Multivariate logistic models indicated that Latino adolescents with an immigrant mother were less likely to report severe psychological distress and that children with an unauthorized father were more likely to report severe psychological distress. These findings reveal important heterogeneity among children in immigrant households and demonstrates the value of measuring legal status in a population survey. It is critical that data used to monitor public health trends more fully incorporate immigrants and their children by measuring domains which are relevant to their health and wellbeing. In addition to measuring what needs to be measured, researchers should continue to critically evaluate quality and put data which are fit to use to meaningful and timely use.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.342 | 0.438 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.006 | 0.011 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".