A Quantitative Analysis on Depictions of Chronic Pain Generated via DALL-E 3, a Text-to-Image Artificial Intelligence Tool
Bibliographic record
Abstract
To the Editor: Chronic pain, pain that persists or recurs for more than 3 months, is a complex health condition affecting approximately 20% of adults.1,2 Visual depictions typically play a pivotal role in how health conditions are perceived and understood; however, this may be more difficult to convey in conditions that are oftentimes “invisible,” such as chronic pain.3 The accurate portrayal of pain states is important not only for clinical assessment but also for raising awareness and cultivating empathy.3,4 Traditional imagery used to represent chronic pain—such as grimacing faces or highlighted body areas—tends to oversimplify the experience, failing to capture its subjective and pervading nature.4 With the advent of artificial intelligence and advances in text-to-image generation models, there is potential to create more nuanced chronic pain representations. Contiguously, the ongoing integration of artificial intelligence into healthcare has the potential to optimize patient flow, diagnostics, and treatment planning.5 However, as artificial intelligence systems become more commonplace, it is crucial to scrutinize data they are trained on and the outputs they generate.6 This Research Letter summarizes a quantitative analysis on depictions of chronic pain generated via OpenAI’s (USA) DALL-E 3, a text-to-image artificial intelligence generation model.7 The aim was to identify inherent biases, explore their origins in the training data, and discuss the implications for the evolving integration of artificial intelligence in healthcare as it pertains to chronic pain. Images (N = 4,000) were generated using a prompt within DALL-E 3. The specific prompt was, “Person with chronic pain (realistic).” A more general prompt was used to examine how artificial intelligence interprets chronic pain without introducing confounding variables, such as detailed descriptors or platform-specific instructions. The generative process was conducted over 20 sessions (1 session per day, 200 images per session), ensuring a variety of outputs and accounting for the artificial intelligence’s stochastic nature. Two authors (M.K. and S.Z.) independently reviewed images for relevance and quality, resolving discrepancies with a third reviewer (E.P.). Inclusion criteria required that images (1) depicted at least one person, (2) were visually coherent without significant distortions or artifacts, and (3) reflected the theme of chronic pain as per the prompt. Images failing to meet these criteria were excluded. Qualifying images were included without further filtering and were analyzed by examining the following: sex, age, race, body habitus, facial affect, physical setting, the presence of medical devices, and the presence of other people (fig. 1). For simplicity and consistency, each variable was described using a binary. Although the methodology for examining each variable was based on an objective, composite analysis of visual characteristics, and contextual cues, the authors acknowledge the possibility of image misinterpretation.Fig. 1.: An example of a text-to-image artificial intelligence–generated depiction of chronic pain on OpenAI’s (USA) DALL-E 3 model using the prompt, “Person with chronic pain (realistic).” The above image would be characterized as: female sex, older age, non-White race, smaller body habitus, negative facial affect, inside setting, present medical devices, and no presence of others.Table 1 presents the results of the quantitative analysis. All images met inclusion criteria. There was a high level of concordance observed among independent reviewers. DALL-E 3’s images appear to follow demographic trends in chronic pain, such as higher prevalence in women and older adults as recently described by Rikard et al. in a review of National Health Interview Survey (National Center for Health Statistics, Hyattsville, Maryland) data, suggesting the potential for artificial intelligence systems to incorporate epidemiologic information effectively.1 However, artificial intelligence’s overexaggeration of trends (e.g., 77.8% images depicted women) risks perpetuating stereotypes and bias into both clinical practice and public perception, while its underrepresentation of certain groups could result in neglect.1,2,8 Of particular note was the limited depiction (5.2%) of individuals with a larger body habitus, despite studies showing strong associations between chronic pain and elevated body mass index.9,10 Table 1. - Results of Quantitative Analysis Chronic Pain Depictions (N = 4,000) M.K. S.Z. Consensus Sex, n (%) Male 888 (22.2) 888 (22.2) 888 (22.2) Female 3,112 (77.8) 3,112 (77.8) 3,112 (77.8) Age, n (%) Younger 1,379 (34.5) 1,364 (34.1) 1,373 (34.3) Older 2,621 (65.5) 2,636 (65.9) 2,627 (65.7) Race, n (%) White 2,224 (55.6) 2,220 (55.5) 2,222 (55.6) Non-White Body habitus, n (%) 1,776 (44.4) 1,780 (44.5) 1,778 (44.4) Smaller 3,793 (94.8) 3,795 (94.9) 3,793 (94.8) Larger 207 (5.2) 205 (5.1) 207 (5.2) Facial affect, n (%) Positive (i.e., smiling) 1,264 (31.6) 1,268 (31.7) 1,265 (31.6) Negative (i.e., frowning) 2,736 (68.4) 2,732 (68.3) 2,735 (68.4) Setting, n (%) Inside 2,836 (70.9) 2,836 (70.9) 2,836 (70.9) Outside 1,164 (29.1) 1,164 (29.1) 1,164 (29.1) Medical devices (present), n (%) Yes 2,012 (50.3) 1,995 (49.9) 2,009 (50.2) No 1,988 (49.7) 2,005 (50.1) 1,991 (49.8) Other people (present), n (%) Yes 1,248 (31.2) 1,248 (31.2) 1,248 (31.2) No 2,752 (68.8) 2,752 (68.8) 2,752 (68.8) Results of analysis performed by the authors (indicated by initials) on 4,000 depictions of chronic pain when the prompt, “Person with chronic pain (realistic)” was inputted into OpenAI’s (USA) DALL-E 3, a text-to-image generator. Both negative facial affect and medical devices (heating/cooling pads, braces/casts, mobility aids, rehabilitation equipment, intravenous therapy) were frequently depicted. The use of negative facial affect and medical devices in artificial intelligence–generated depictions of chronic pain highlights a tension in visually representing an oftentimes “invisible” condition. Chronic pain frequently lacks outward physical manifestations, making it difficult to communicate or validate using visual cues alone. These two elements—facial affect and medical devices—serve as visual shorthand to convey the presence of chronic pain, but also raise critical questions about how chronic pain is perceived and understood. Overemphasis on negative facial affect and medical devices risks reinforcing a narrow view that chronic pain must always “look” painful to be legitimate. This may exacerbate the already significant challenges that patients face in having chronic pain acknowledged and treated seriously. Relatedly, most images depicted indoor settings and individuals experiencing chronic pain in isolation. This reflects a common narrative of chronic pain as a solitary struggle, which, while aligning with the reality that chronic pain may lead to social withdrawal or isolation, neglects to demonstrate the importance of relationships and social support systems in chronic pain management. It remains indeterminate what combination of input data are driving these findings. Results may be due to various factors, including cultural perceptions that some conditions are more common in specific groups; stereotypes about expressiveness or vulnerability affecting perceptions and reporting; studies indicating higher reported rates of certain conditions within groups; increased visibility in clinical datasets due to a group’s likelihood of seeking medical care for related issues; disproportionate portrayal in media or artistic sources skewing public perception; datasets that focus on specific groups leading to overrepresentation; underreporting or understudying of conditions in other groups; and artificial intelligence systems trained on unbalanced datasets that reflect and amplify these disparities. Text-to-image models, while seemingly rudimentary, transform training data into visual outputs, enabling the identification and analysis of patterns with clarity and efficiency. Analyzing how models depict concepts like pain or patient demographics can highlight embedded concerns. While DALL-E 3’s images capture some key demographic trends in chronic pain, they also potentially reinforce biases present in training datasets. A concerted effort to diversify data, refine algorithms, and align artificial intelligence outputs with real-world demographics is essential for improving the effectiveness of artificial intelligence tools intended for use in healthcare and chronic pain management, for both clinicians and patients alike. In the interim, clinicians should exercise caution in the use, and consideration, of artificial intelligence–generated content. Further investigation in this area may involve more nuanced image analysis (e.g., precise demographic characterization and setting interpretation), inquiries into relationships between variables (e.g., sex and age), examination of how different chronic pain states are depicted (e.g., fibromyalgia vs. complex regional pain syndrome), exploration of other models (e.g., Midjourney), and descriptions of visual art elements (e.g., color and lighting). Competing Interests The authors declare no competing interests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".