MétaCan
Menu
Retour à la cohorte
Enregistrement W4407366799 · doi:10.1097/aln.0000000000005364

A Quantitative Analysis on Depictions of Chronic Pain Generated via DALL-E 3, a Text-to-Image Artificial Intelligence Tool

2025· article· en· W4407366799 sur OpenAlexaff
Morgan King

Notice bibliographique

RevueAnesthesiology · 2025
Typearticle
Langueen
DomaineMedicine
ThématiqueTraditional Chinese Medicine Studies
Établissements canadiensUniversity of Toronto
Organismes subventionnairesnon disponible
Mots-clésMedicineArtificial intelligence

Résumé

récupéré en direct d'OpenAlex

To the Editor: Chronic pain, pain that persists or recurs for more than 3 months, is a complex health condition affecting approximately 20% of adults.1,2 Visual depictions typically play a pivotal role in how health conditions are perceived and understood; however, this may be more difficult to convey in conditions that are oftentimes “invisible,” such as chronic pain.3 The accurate portrayal of pain states is important not only for clinical assessment but also for raising awareness and cultivating empathy.3,4 Traditional imagery used to represent chronic pain—such as grimacing faces or highlighted body areas—tends to oversimplify the experience, failing to capture its subjective and pervading nature.4 With the advent of artificial intelligence and advances in text-to-image generation models, there is potential to create more nuanced chronic pain representations. Contiguously, the ongoing integration of artificial intelligence into healthcare has the potential to optimize patient flow, diagnostics, and treatment planning.5 However, as artificial intelligence systems become more commonplace, it is crucial to scrutinize data they are trained on and the outputs they generate.6 This Research Letter summarizes a quantitative analysis on depictions of chronic pain generated via OpenAI’s (USA) DALL-E 3, a text-to-image artificial intelligence generation model.7 The aim was to identify inherent biases, explore their origins in the training data, and discuss the implications for the evolving integration of artificial intelligence in healthcare as it pertains to chronic pain. Images (N = 4,000) were generated using a prompt within DALL-E 3. The specific prompt was, “Person with chronic pain (realistic).” A more general prompt was used to examine how artificial intelligence interprets chronic pain without introducing confounding variables, such as detailed descriptors or platform-specific instructions. The generative process was conducted over 20 sessions (1 session per day, 200 images per session), ensuring a variety of outputs and accounting for the artificial intelligence’s stochastic nature. Two authors (M.K. and S.Z.) independently reviewed images for relevance and quality, resolving discrepancies with a third reviewer (E.P.). Inclusion criteria required that images (1) depicted at least one person, (2) were visually coherent without significant distortions or artifacts, and (3) reflected the theme of chronic pain as per the prompt. Images failing to meet these criteria were excluded. Qualifying images were included without further filtering and were analyzed by examining the following: sex, age, race, body habitus, facial affect, physical setting, the presence of medical devices, and the presence of other people (fig. 1). For simplicity and consistency, each variable was described using a binary. Although the methodology for examining each variable was based on an objective, composite analysis of visual characteristics, and contextual cues, the authors acknowledge the possibility of image misinterpretation.Fig. 1.: An example of a text-to-image artificial intelligence–generated depiction of chronic pain on OpenAI’s (USA) DALL-E 3 model using the prompt, “Person with chronic pain (realistic).” The above image would be characterized as: female sex, older age, non-White race, smaller body habitus, negative facial affect, inside setting, present medical devices, and no presence of others.Table 1 presents the results of the quantitative analysis. All images met inclusion criteria. There was a high level of concordance observed among independent reviewers. DALL-E 3’s images appear to follow demographic trends in chronic pain, such as higher prevalence in women and older adults as recently described by Rikard et al. in a review of National Health Interview Survey (National Center for Health Statistics, Hyattsville, Maryland) data, suggesting the potential for artificial intelligence systems to incorporate epidemiologic information effectively.1 However, artificial intelligence’s overexaggeration of trends (e.g., 77.8% images depicted women) risks perpetuating stereotypes and bias into both clinical practice and public perception, while its underrepresentation of certain groups could result in neglect.1,2,8 Of particular note was the limited depiction (5.2%) of individuals with a larger body habitus, despite studies showing strong associations between chronic pain and elevated body mass index.9,10 Table 1. - Results of Quantitative Analysis Chronic Pain Depictions (N = 4,000) M.K. S.Z. Consensus Sex, n (%) Male 888 (22.2) 888 (22.2) 888 (22.2) Female 3,112 (77.8) 3,112 (77.8) 3,112 (77.8) Age, n (%) Younger 1,379 (34.5) 1,364 (34.1) 1,373 (34.3) Older 2,621 (65.5) 2,636 (65.9) 2,627 (65.7) Race, n (%) White 2,224 (55.6) 2,220 (55.5) 2,222 (55.6) Non-White Body habitus, n (%) 1,776 (44.4) 1,780 (44.5) 1,778 (44.4) Smaller 3,793 (94.8) 3,795 (94.9) 3,793 (94.8) Larger 207 (5.2) 205 (5.1) 207 (5.2) Facial affect, n (%) Positive (i.e., smiling) 1,264 (31.6) 1,268 (31.7) 1,265 (31.6) Negative (i.e., frowning) 2,736 (68.4) 2,732 (68.3) 2,735 (68.4) Setting, n (%) Inside 2,836 (70.9) 2,836 (70.9) 2,836 (70.9) Outside 1,164 (29.1) 1,164 (29.1) 1,164 (29.1) Medical devices (present), n (%) Yes 2,012 (50.3) 1,995 (49.9) 2,009 (50.2) No 1,988 (49.7) 2,005 (50.1) 1,991 (49.8) Other people (present), n (%) Yes 1,248 (31.2) 1,248 (31.2) 1,248 (31.2) No 2,752 (68.8) 2,752 (68.8) 2,752 (68.8) Results of analysis performed by the authors (indicated by initials) on 4,000 depictions of chronic pain when the prompt, “Person with chronic pain (realistic)” was inputted into OpenAI’s (USA) DALL-E 3, a text-to-image generator. Both negative facial affect and medical devices (heating/cooling pads, braces/casts, mobility aids, rehabilitation equipment, intravenous therapy) were frequently depicted. The use of negative facial affect and medical devices in artificial intelligence–generated depictions of chronic pain highlights a tension in visually representing an oftentimes “invisible” condition. Chronic pain frequently lacks outward physical manifestations, making it difficult to communicate or validate using visual cues alone. These two elements—facial affect and medical devices—serve as visual shorthand to convey the presence of chronic pain, but also raise critical questions about how chronic pain is perceived and understood. Overemphasis on negative facial affect and medical devices risks reinforcing a narrow view that chronic pain must always “look” painful to be legitimate. This may exacerbate the already significant challenges that patients face in having chronic pain acknowledged and treated seriously. Relatedly, most images depicted indoor settings and individuals experiencing chronic pain in isolation. This reflects a common narrative of chronic pain as a solitary struggle, which, while aligning with the reality that chronic pain may lead to social withdrawal or isolation, neglects to demonstrate the importance of relationships and social support systems in chronic pain management. It remains indeterminate what combination of input data are driving these findings. Results may be due to various factors, including cultural perceptions that some conditions are more common in specific groups; stereotypes about expressiveness or vulnerability affecting perceptions and reporting; studies indicating higher reported rates of certain conditions within groups; increased visibility in clinical datasets due to a group’s likelihood of seeking medical care for related issues; disproportionate portrayal in media or artistic sources skewing public perception; datasets that focus on specific groups leading to overrepresentation; underreporting or understudying of conditions in other groups; and artificial intelligence systems trained on unbalanced datasets that reflect and amplify these disparities. Text-to-image models, while seemingly rudimentary, transform training data into visual outputs, enabling the identification and analysis of patterns with clarity and efficiency. Analyzing how models depict concepts like pain or patient demographics can highlight embedded concerns. While DALL-E 3’s images capture some key demographic trends in chronic pain, they also potentially reinforce biases present in training datasets. A concerted effort to diversify data, refine algorithms, and align artificial intelligence outputs with real-world demographics is essential for improving the effectiveness of artificial intelligence tools intended for use in healthcare and chronic pain management, for both clinicians and patients alike. In the interim, clinicians should exercise caution in the use, and consideration, of artificial intelligence–generated content. Further investigation in this area may involve more nuanced image analysis (e.g., precise demographic characterization and setting interpretation), inquiries into relationships between variables (e.g., sex and age), examination of how different chronic pain states are depicted (e.g., fibromyalgia vs. complex regional pain syndrome), exploration of other models (e.g., Midjourney), and descriptions of visual art elements (e.g., color and lighting). Competing Interests The authors declare no competing interests.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,001
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,752
Score d'incertitude au seuil0,538

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,001
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0010,000
Bibliométrie0,0010,002
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,034
Tête enseignante GPT0,335
Écart entre enseignants0,301 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueAnesthesiologyMême sujetTraditional Chinese Medicine StudiesTravaux en français237 207