A Pilot Study of Electronic Cardiovascular Operative Notes: Qualitative Assessment and Challenges in Implementation
Bibliographic record
Abstract
BACKGROUND: Our objectives are to describe the contents of cardiovascular surgical operative notes and to develop and test a standards-based structured electronic operative note that might be used for secondary purposes. STUDY DESIGN: Operative notes were selected for patients who underwent primary, isolated coronary artery bypass grafting (n = 33); aortic valve replacement (n = 33); reoperative coronary artery bypass grafting (n = 11); or aortic valve replacement (n = 11). The content was qualitatively assessed and categorized into 3 sections, ie, technical/procedural, anatomic/physiologic description, and judgment/opinion. An electronic operative note was developed using a standards-based approach to categorize the type of operation. RESULTS: Average length +/- SD of the operative note was 495 +/- 186 words (range 243 to 1,267 words). The procedural category made up a mean proportion of 73% +/- 12% (range 32% to 95%). The descriptive category was the second largest category in the operative note; mean percentage 22% +/- 8% (range 5% to 43%). The dictation of the judgment portion made up 6% +/- 6% (range 0% to 25%) of the operative note. In the pilot electronic note system, 5 surgeons entered 23 procedures performed on 18 patients (14% of eligible patients). Seventeen (74%) procedures entered by surgeons were in complete agreement with the data for the Society of Thoracic Surgeons database collected by professional abstractors. CONCLUSIONS: Freeform dictation of cardiovascular notes varied by individual surgeon style and case complexity. Up to 25% of the operative note was dedicated to judgment/opinion, which would be difficult to recreate in a structured data-entry format. An electronic system for entering procedural details can improve efficiency for secondary purposes of data collection but must be carefully implemented to avoid loss of important information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".