Development and Validation of an Algorithm for Quality Grading of Pediatric Spirometry: A Quality Improvement Initiative
Bibliographic record
Abstract
Abstract Rationale Current spirometry quality grading for individuals 7 years and older include within-test repeatability thresholds up to 250 ml, which may be inappropriately wide for children. Objectives 1) To develop, internally validate, and implement a quality grading algorithm for forced expiratory volume in 1 second (FEV1) and forced vital capacity (FVC) for school-aged children and 2) to compare the algorithm with the one proposed by the American Thoracic Society (ATS). Methods We conducted a review of existing algorithms and obtained expert input. A pediatric quality grading algorithm was drafted and modified in an iterative process until consensus was achieved, with the main difference from current criteria being tighter volume repeatability for the pediatric quality grading. Four pulmonary function technicians evaluated the interrater agreement of the algorithm in a blinded fashion on an unselected consecutive sample of 87 prebronchodilator spirometry tests, and the grades were compared with those from the ATS algorithm in the same sample of spirometry tests. The algorithm was then implemented into the workflow of the pulmonary function laboratory. Results For FEV1 and FVC, the interrater agreement values for the pediatric algorithm were 92% and 83%, respectively. When the ATS algorithm was used, 75.9% (n = 66) and 63.2% (n = 55) of subjects achieved a grade of A for FEV1 and FVC; when the pediatric algorithm was used, 69.0% (n = 60) and 58.6% (n = 51) met grade A criteria. There was a more uniform distribution of tests for the pediatric algorithm across grades B through F for FEV1 and FVC, and no grade C tests were observed for the ATS grading algorithm. A total of 2,464 tests graded prospectively by using the pediatric algorithm showed a median (interquartile range) FEV1 and FVC repeatability within 29 ml (13–57 ml) and 34 ml (15–66 ml), respectively. Most subjects received a grade of A for FEV1 (81.1%) and FVC (71.6%), performing a repeatable spirometry test to within 100 ml. Conclusions A quality grading algorithm that uses smaller ranges of expired volumes to define repeatability is feasible and may be more appropriate in a pediatric pulmonary function laboratory.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.209 | 0.202 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.010 | 0.005 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".