Population Health across Space and Time: the Geographical Harmonisation of the Office for National Statistics Longitudinal Study for England and Wales
Bibliographic record
Abstract
ABSTRACT There is a need in health research to identify whether inequalities are increasing or improving between different geographical areas. Both cross‐sectional time‐series and longitudinal/cohort studies contribute to our knowledge, with the Office for National Statistics Longitudinal Study (LS) for England and Wales being a major resource. However, any research into geographical change over time can be hampered by boundary change or when the geographical definition for which data are available is not the geography relevant to an analysis. We develop a method using population‐weighted centroids of estimating an LS member's location at a previous time point and then link this to a small‐area geography, the 2001 Census Output Areas. This is not so that analyses can be carried out at this scale but so that records can be linked to larger geographies or area classifications. A time‐series or longitudinal analysis can then be carried out and geographical trends observed. In terms of reliability, we find that accuracy improves with increasing size of geographical units and when area typologies are used. In example analyses using a geodemographic classification attached to LS members' records, we find that in a time series of cross‐sections, mortality improves across all area types but not to the same extent. A longitudinal analysis indicates that changes in the area types in which people were living lead to steeper health gradients than if people had stayed living in the same type of area. Differences, though, are small, suggesting that, in the main, there is little mobility between area types. We recommend that longitudinal and cohort studies retain the postcode of each member's address so that ongoing linkages can be made when administrative boundary changes occur and for relevance to application relevant geographies. Our method can be used to enhance previous records and thereby maximise previous investment in the collection of data. Copyright © 2012 John Wiley & Sons, Ltd.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.039 | 0.053 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.010 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.005 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".