Challenges and benefits of integrating diverse sampling strategies in the observation of cardiovascular risk factors (ORISCAV-LUX 2) study
Bibliographic record
Abstract
BACKGROUND: It is challenging to manage data collection as planned and creation of opportunities to adapt during the course of enrolment may be needed. This paper aims to summarize the different sampling strategies adopted in the second wave of Observation of Cardiovascular Risk Factors (ORISCAV-LUX, 2016-17), with a focus on population coverage and sample representativeness. METHODS: Data from the first nationwide cross-sectional, population-based ORISCAV-LUX survey, 2007-08 and from the newly complementary sample recruited via different pathways, nine years later were analysed. First, we compare the socio-demographic characteristics and health profiles between baseline participants and non-participants to the second wave. Then, we describe the distribution of subjects across different strategy-specific samples and performed a comparison of the overall ORISCAV-LUX2 sample to the national population according to stratification criteria. RESULTS: For the baseline sample (1209 subjects), the participants (660) were younger than the non-participants (549), with a significant difference in average ages (44 vs 45.8 years; P = 0.019). There was a significant difference in terms of education level (P < 0.0001), 218 (33%) participants having university qualification vs. 95 (18%) non-participants. The participants seemed having better health perception (p < 0.0001); 455 (70.3%) self-reported good or very good health perception compared to 312 (58.2%) non-participants. The prevalence of obesity (P < 0.0001), hypertension (P < 0.0001), diabetes (P = 0.007), and mean values of related biomarkers were significantly higher among the non-participants. The overall sample (1558 participants) was mainly composed of randomly selected subjects, including 660 from the baseline sample and 455 from other health examination survey sample and 269 from civil registry sample (constituting in total 88.8%), against only 174 volunteers (11.2%), with significantly different characteristics and health status. The ORISCAV-LUX2 sample was representative of national population for geographical district, but not for sex and age; the younger (25-34 years) and older (65-79 years) being underrepresented, whereas middle-aged adults being over-represented, with significant sex-specific difference (p < 0.0001). CONCLUSION: This study represents a careful first-stage analysis of the ORISCAV-LUX2 sample, based on available information on participants and non-participants. The ORISCAV-LUX datasets represents a relevant tool for epidemiological research and a basis for health monitoring and evidence-based prevention of cardiometabolic risk in Luxembourg.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.516 | 0.433 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.008 | 0.011 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".