Replication Data for: The impact of physical activity cut-point choice on childhood activity estimates
Bibliographic record
Abstract
This dataset is the replication data for the paper "The impact of physical activity cut-point choice on childhood activity estimates". It includes accelerometer data from 617 children who contributed a total of 1314 participant-weeks of valid data after processing. The detailed description of each column is as follows: 1. participant_week_index. Because we are considering participant week as our basic analysis object, so each participant week is assigned with a unique index. 2. datetime. The time is at the minute level. 3. time_class. There are 3 types of time in our study which are: 3.1. time class [1] represents school time. 09:15 am to 15:00 pm from Monday to Friday. 3.2. time class [2] represents leisure time. 06:00am to 09:15am and 15:00pm to 22:00pm from Monday to Friday, and 06:00am to 22:00pm on Saturday and Sunday. 3.3. time class [3] represents other time. 4. counts_per_minute. The accelerometer data were collected at 100 Hz epochs and reduced to vector magnitude (VM) with a 1-second epoch using ActiLife 6 data analysis software. Further, the VM within a minute is accumulated to get counts_per_minute. 5. wear_and_awake. There are 2 values: 5.1. value [0] represents that the participant is sleeping or doesn't wear the device. 5.2. value [1] represents the participant is awake and wear the device. 6. activity_level. There are 4 types under standard threshold: 6.1. N/A represents "not applicable". Because we only consider the accelerometer data during leisure hours, the activity level will be labeled as N/A if the time is not in leisure hours. 6.2. SED if CPM <= 150. 6.3. LPA if CPM is in (150, 1951] 6.4. MVPA if CPM > 1951
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Dataset About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
| gpt | no category Domain: not available · Genre: Dataset About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.052 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.046 | 0.025 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".