Validität und Reliabilität eines Instruments zur Messung der Qualität der Kommunikation und seine Eignung im studentischen Unterricht
Bibliographic record
Abstract
Questions and aims: The aim of this study (thesis) is to investigate if a short-version of the „Calgary-Cambridge Observation Guide for the Medical Interview“ (CCOG) by Kurtz and Silverman (1996), translated into German, is valid and reliable and can be used to ojectively judge the communicational abilities of medical students. Method: In intervals of at least three months a selected group of physicians, re-search assistants and medical students evaluated five videos of anamnesis using the short-version German translation of the CCOG. Each video detailed a different level of communicative skills on the part of the student playing the medical doctor. The evaluation was made based on the following criteria: time of assessment, sex, group of evaluators, quality of video conversation. Moreover, an explorative and confirmatory factor analysis was calculated and the retest reliability as well as the intra-class-correlation was determined. Results: 30 evaluators took part in the study, 3 of which as so-called ‘gold standard’. The evaluation of all 5 videos of anamnesis showed a slight improvement of marks at the second evaluation time. In the original version the CCOG contains 28 items, summed up in 6 scales of 3 to 7 items each. This structure could only partly be de-picted in the performed factor analysis. According to value (“eigenvalue”) > 1 five factors were sufficient to illustrate and divide the items. Moreover, a different scale allocation from the original was shown, allocating more than half of the items (15/28) to the same factor. Also, the inter-rater agreement in reference to answering single items was not optimal (ICC range 0.05 – 0.57). Conclusions: The short version of the CCOG showed a good agreement at the re-test-reliability. Considerable differences were observed in the evaluation of some items when comparing between the ‘gold standard’ and the other evaluators. The structure of the item scales and the inter-rater reliability are only conditionally ac-ceptable. Perhaps a 3-stage evaluation scale, a more homogeneous group of evaluators or a better training might have improved the result of the ICC. Some items should be discarded, rephrased or combined in a better way or new items should be added and scales be restructured.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.084 | 0.108 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".