Evaluating the accuracy and consistency of ChatGPT for the management of type 2 diabetes: A cross-sectional study
Bibliographic record
Abstract
Abstract Large language models (LLMs) have fundamentally changed how patients and clinicians retrieve information; however, it is unclear how accurate and consistent widely available LLMs are in answering questions related to medical information. Our objective was to evaluate the accuracy and consistency of ChatGPT in answering questions related to the management of type 2 diabetes mellitus (T2DM). Three users asked ChatGPT 13 questions pertaining to medications from the top five most common classes of T2DM medications. A response was labelled inconsistent if the response provided to one user differed from the response provided to at least one other user in the same domain for the same medication. A response was labelled as inaccurate if the information provided by ChatGPT was incorrect based on the most recent FDA-approved drug label, in addition to review by an expert reviewer. Additionally, one user asked ChatGPT 26 basic questions related to the management of T2DM, in which the answer was categorized as correct or incorrect. We summarized all results using descriptive statistics. ChatGPT delivered inaccurate responses in seven out of 13 domains and inconsistent responses in seven out of 13 domains for drugs in all five classes of T2DM medication. Of ChatGPT’s responses to the 26 basic T2DM treatment questions, 7 (26%) were incorrect. In this cross-sectional study, we identified that it was common for ChatGPT to provide incorrect or inconsistent responses to enquiries related to the management of type 2 diabetes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.050 | 0.177 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".