Construction and validation of the area level deprivation index for health research: A methodological study based on Nepal Demographic and Health Survey
Bibliographic record
Abstract
Area-level factors may partly explain the heterogeneity in risk factors and disease distribution. Yet, there are a limited number of studies that focus on the development and validation of the area level construct and are primarily from high-income countries. The main objective of the study is to provide a methodological approach to construct and validate the area level construct, the Area Level Deprivation Index in low resource setting. A total of 14652 individuals from 11,203 households within 383 clusters (or areas) were selected from 2016-Nepal Demographic and Health survey. The index development involved sequential steps that included identification and screening of variables, variable reduction and extraction of the factors, and assessment of reliability and validity. Variables that could explain the underlying latent structure of area-level deprivation were selected from the dataset. These variables included: housing structure, household assets, and availability and accessibility of physical infrastructures such as roads, health care facilities, nearby towns, and geographic terrain. Initially, 26-variables were selected for the index development. A unifactorial model with 15-variables had the best fit to represent the underlying structure for area-level deprivation evidencing strong internal consistency (Cronbach's alpha = 0.93). Standardized scores for index ranged from 58.0 to 140.0, with higher scores signifying greater area-level deprivation. The newly constructed index showed relatively strong criterion validity with multi-dimensional poverty index (Pearson's correlation coefficient = 0.77) and relatively strong construct validity (Comparative Fit Index = 0.96; Tucker-Lewis Index = 0.94; standardized root mean square residual = 0.05; Root mean square error of approximation = 0.079). The factor structure was relatively consistent across different administrative regions. Area level deprivation index was constructed, and its validity and reliability was assessed. The index provides an opportunity to explore the area-level influence on disease outcome and health disparity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.055 | 0.078 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".