Canadian Spine Society Abstracts 2017
Bibliographic record
Abstract
Background: The Canadian Spine Outcomes and Research Network (CSORN) is an emergent, rapidly growing national spine registry.The utility of a medical database is dependent on the quality of the data.Our objective is to evaluate data quality of the CSORN database.Methods: A retrospective analysis of data from 17 CSORN sites across Canada, including 6233 patients, was conducted.Completeness of data, follow-up rates and percentage of patient enrolment were assessed via CSORN data quality reports and site interviews.Data quality was operationally defined as poor (< 60% complete), moderate (60%-80% complete) and good (> 80% complete).Descriptive statistics were used to ascertain data completeness and follow-up rates.Repeated-measures analysis of variance was used to investigate the effect of time.Significance was set at α < 0.05.Results: Analy ses revealed successful enrolment of 78% of potential patients.Follow-up rates significantly decreased across time (F 1,30 = 15.10,p = 0.001).Follow-up rates declined significantly from 12 weeks (87.76%) to 12 months (59.25%) and 24 months (44.68%).At 12 weeks, 82.35% of sites had good follow-up, and 17.6% were moderate.At 12 and 24 months, 25% of sites had good follow-up rates, 25% had moderate and 50% had poor follow-up, as defined by this study.Overall data quality averaged 86.14%.The primary issue with data completeness can be narrowed down to specific problematic variables.The 24-month data quality is significantly lower, with thoracolumbar and cervical follow-up at moderate (72.65%) and poor (53.10%) quality, respectively.Mapping data from a previous database resulted in limited data quality.On average, 42.85% of CSORN variables were not included in previous databases, resulting in a 22.64% data quality drop.Conclusion: Overall data quality was classified as good.For new sites considering integration into CSORN, we do not recommend mapping over previously collected data.However, follow-up rates at 12 and 24 months were poor.This is common in medical databases, and CSORN has taken steps to improve data quality, including hiring a data quality coordinator and new training.Data quality is a critical component of databases and reanalysis is warranted following planned interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.004 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.005 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.610 | 0.448 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".