A-180 Counting CAP deficiency Caused by Clerical errors: a regional study
Bibliographic record
Abstract
Abstract Background Proficiency testing (PT) is a critical component of quality assurance and a laboratory accreditation requirement. PT ensures that laboratory results are comparable to those of peers using the same instruments and methods. Failures may indicate inaccurate patient results, requiring formal investigation and corrective action by the testing laboratory. Depending on the accreditation body, consecutive PT deficiencies can lead to mandatory cessation of testing. Clerical errors are among the most common causes of unacceptable PT grades. While many laboratory instruments directly communicate with the laboratory information system for electronic result transfer, PT result reporting often involves manual transcription. This study aims to determine the frequency of clerical errors associated with unacceptable PT grades and to implement quality improvement strategies where necessary. Methods All PT results submitted to the College of American Pathologists (CAP) between January 2020 and September 2024 (57 months) were reviewed for three chemistry laboratories (designated as A, B, and C) in Saskatoon, Saskatchewan, Canada. The number of unacceptable CAP submissions was tabulated and compared across the three laboratories. Clerical errors contributing to unacceptable submissions were classified into three subcategories: direct transcription errors, incorrect method codes, and incorrect instrument codes. Chi-squared analysis was performed to assess statistical significance. The CAP PT reporting process in the core chemistry laboratory was reviewed through discussions with senior technologists responsible for result submission. Results A total of 32,868 results were submitted to CAP during the study period: Laboratory A (16,456), Laboratory B (2,688), and Laboratory C (13,724). The overall percentage of unacceptable CAP submissions was 0.83% (273/32,868). Of these, 44% (120/273) were attributed to clerical errors, including direct transcription errors (43/120), incorrect method codes (49/120), and incorrect instrument codes (27/120). Laboratory A had a significantly lower percentage of clerical errors [0.18% (30/16,456)] compared to Laboratory B [0.56% (15/2,688)] and Laboratory C [0.55% (75/13,724)] (p < 0.001, Chi-squared = 30, df = 2). No significant difference was observed between Laboratories B and C. A process review revealed that Laboratory A was the only site utilizing a secondary verification step, where a second technologist reviewed PT results before submission. Conclusion Clerical errors are a major contributor to CAP deficiencies. Laboratory A, which implements a secondary verification process for manually entered PT results, had significantly fewer clerical errors. This practice aligns with established procedures for manual patient result entry. Consequently, Laboratories B and C have now adopted a similar verification process. Follow-up studies will evaluate whether this intervention reduces clerical errors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".