Interpretation of Computed Tomography of the Head: Emergency Physicians versus Radiologists
Bibliographic record
Abstract
BACKGROUND: Many patients are brought to crowded emergency departments (ED) of hospitals every day for evaluation of head injuries, headaches, neurologic deficits etc. CT scan of the head is the most common diagnostic measure used to search for pathologies. In many EDs the initial interpretation of images are performed by emergency physicians (EP). Since most decisions are made based on the initial interpretation of the images by emergency physicians and not the radiologists, it is necessary to assess the accuracy of interpretations made by the former group. OBJECTIVES: The objective of this study was to compare the findings reported in the interpretation of head CTs by emergency physicians and compare to radiologists (the gold standard). MATERIALS AND METHODS: This was a prospective cross sectional study conducted from March to May 2009 in a teaching hospital in Tehran, Iran. All non-contrast head CTs obtained during the study period were copied on DVDs and sent separately to a radiologist, 6 emergency medicine (EM) attending physicians and 14 senior EM residents for interpretation. Clinical information pertaining to each patient was also sent with each CT. The radiologist's interpretation was considered as the gold standard and reference for comparison. Data from EM physicians and residents were compared with the reference as well as with each other and statistical analysis was performed using SPSS 18.5. RESULTS: Out of 544 CT scans, EM physicians had 35 false negatives and 53 false positives compared with radiologist's interpretations (P < 0.0001). EM residents had 74 false negatives and 12 false positives compared with radiologist's interpretations (P < 0.0001). CONCLUSIONS: Both EPs and ER residents either missed or falsely called a significant number of pathologies in their interpretations. The interpretations of EPs and ER residents were more sensitive and more specific, respectively. These findings revealed the need for increased training time in head CT reading for residents and the necessity of attending continuing medical education workshops for emergency physicians.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.048 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".