Is the lack of smartphone data skewing wealth indices in low-income settings?
Bibliographic record
Abstract
BACKGROUND: Smartphones have rapidly become an important marker of wealth in low- and middle-income countries, but international household surveys do not regularly gather data on smartphone ownership and these data are rarely used to calculate wealth indices. METHODS: We developed a cross-sectional survey module delivered to 3028 households in rural northwest Burkina Faso to measure the effects of this absence. Wealth indices were calculated using both principal components analysis (PCA) and polychoric PCA for a base model using only ownership of any cell phone, and a full model using data on smartphone ownership, the number of cell phones, and the purchase of mobile data. Four outcomes (household expenditure, education level, and prevalence of frailty and diabetes) were used to evaluate changes in the composition of wealth index quintiles using ordinary least squares and logistic regressions and Wald tests. RESULTS: Households that own smartphones have higher monthly expenditures and own a greater quantity and quality of household assets. Expenditure and education levels are significantly higher at the fifth (richest) socioeconomic status (SES) quintile of full model wealth indices as compared to base models. Similarly, diabetes prevalence is significantly higher at the fifth SES quintile using PCA wealth index full models, but this is not observed for frailty prevalence, which is more prevalent among lower SES households. These effects are not present when using polychoric PCA, suggesting that this method provides additional robustness to missing asset data to measure underlying latent SES by proxy. CONCLUSIONS: The lack of smartphone data can skew PCA-based wealth index performance in a low-income context for the top of the socioeconomic spectrum. While some PCA variants may be robust to the omission of smartphone ownership, eliciting smartphone ownership data in household surveys is likely to substantially improve the validity and utility of wealth estimates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.067 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".