Reliability of routinely collected anthropometric measurements in primary care
Bibliographic record
Abstract
BACKGROUND: Measuring body mass index (BMI) has been proposed as a method of screening for preventive primary care and population surveillance of childhood obesity. However, the accuracy of routinely collected measurements has been questioned. The purpose of this study was to assess the reliability of height, length and weight measurements collected during well-child visits in primary care relative to trained research personnel. METHODS: A cross-sectional study of measurement reliability was conducted in community pediatric and family medicine primary care practices. Each participating child, ages 0 to 18 years, was measured four consecutive times; twice by a primary care team member (e.g. nurses, practice personnel) and twice by a trained research assistant. Inter- and intra-observer reliability was calculated using the technical error of measurement (TEM), relative TEM (%TEM), and a coefficient of reliability (R). RESULTS: Six trained research assistants and 16 primary care team members performed measurements in three practices. All %TEM values for intra-observer reliability of length, height, and weight were classified as 'acceptable' (< 2%; range 0.19% to 0.70%). Inter-observer reliability was also classified as 'acceptable' (< 2%; range 0.36% to 1.03%) for all measurements. Coefficients of reliability (R) were all > 99% for both intra- and inter-observer reliability. Length measurements in children < 2 years had the highest measurement error. There were some significant differences in length intra-observer reliability between observers. CONCLUSION: There was agreement between routine measurements and research measurements although there were some differences in length measurement reliability between practice staff and research assistants. These results provide justification for using routinely collected data from selected primary care practices for secondary purposes such as BMI population surveillance and research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.120 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".