Reference equations for pulmonary diffusing capacity using segmented regression show similar predictive accuracy as GAMLSS models
Bibliographic record
Abstract
PURPOSE: To determine whether generalised additive models of location, scale and shape (GAMLSS) developed for pulmonary diffusing capacity are superior to segmented (piecewise) regression models, and to update reference equations for pulmonary diffusing capacity for carbon monoxide (DLCO) and nitric oxide (DLNO), which may be affected by the equipment used for its measurement. METHODS: ). Reference equations were created for DLCO and DLNO using both GAMLSS and segmented linear regression. Cross-validation was applied to compare the prediction accuracy of the two models as follows: 80% of the pooled data were used to create the equations, and the remaining 20% was used to examine the fit. This was repeated 100 times. Then, the root-mean-square error was compared between both models. RESULTS: In males, GAMLSS models were 7% worse to 3% better compared to segmented regression for DLCO and DLNO. In females, GAMLSS models were 2% worse to 5% better compared to segmented linear regression for DLCO and DLNO. The Hyp'Air Compact measured DLNO and alveolar volume (VA) that was approximately 16-20 mL/min/mm Hg and 0.2-0.4 L higher, respectively, compared to the Jaeger MasterScreen Pro. The measured DLCO was similar between devices after controlling for altitude. CONCLUSIONS: For the development of pulmonary function reference equations, we propose that segmented linear regression can be used instead of GAMLSS due to its simplicity, especially when the predictive accuracy is similar between the two models, overall.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.035 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".