Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

Abstract Background Generative AI video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions such as depression. Because AI-generated content may reach people during vulnerable periods, understanding what visual narratives these models produce for sensitive concepts carries clinical relevance. Objective This study aimed to characterize how OpenAI’s Sora 2 generative AI video model depicts depression and examine whether depictions differ between the consumer app and developer API access points, which differ in their product layer mediation. Methods We generated 100 videos using the single-word prompt “Depression” across 2 access points: the consumer app (n=50, 50%) and developer API (n=50, 50%). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Interrater reliability was assessed using the Cohen κ, with dimensions showing insufficient agreement excluded from analysis. Computational features (visual aesthetics, audio, semantic content, and temporal dynamics) were extracted and compared between modalities using 2-tailed Welch t tests with Benjamini-Hochberg false discovery rate correction. Results App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution compared with 14% (7/50) of API outputs. This divergence was reinforced across channels. App videos brightened over time (mean slope 2.90, SD 2.43 per second vs −0.18, SD 1.24 per second for the API; Cohen d =1.59; q <.001) and contained 3 times more motion (Cohen d =2.07; q <.001). Across both modalities, videos converged on a narrow visual vocabulary: predominantly seated figures (94/100, 94% of the videos); downward gaze (93/100, 93% of the videos); and recurring objects including hoodies (n=194), windows (n=148), and rain (n=83). Transcript language in depressive phases emphasized weight and containment (“heavy,” “drowning,” and “room”), whereas recovery phases reversed these patterns: brightness increased by 26.6% (Cohen d =0.68; P <.001), gaze shifted upward in 67% (30/45) of recovery videos, and terms such as “light” and “breath” emerged. Figures were predominantly young adults (323/367, 88% aged 20-30 years) and nearly always alone (360/367, 98%). Gender varied by access point: app outputs skewed male (34/50, 68%), and API outputs skewed female (33/56, 59%). Conclusions Sora 2 does not invent new visual grammars for depression but compresses and recombines cultural iconographies, whereas platform-level constraints substantially shape which narratives reach users. Clinicians should be aware that AI-generated mental health video content reflects training data and platform design rather than clinical knowledge and that patients may encounter such content during vulnerable periods.

Authors

Publication Details

Journal
Journal of Medical Internet Research
Published
2026-09-14
DOI
https://doi.org/10.2196/95682
Primary Topic
Digital Mental Health Interventions
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

Julian Herpertz, John Torous, Matthew Flathers, Griffin Smith et al.
Journal of Medical Internet Research
Digital Mental Health Interventions
article

Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

Julian Herpertz, John Torous, Matthew Flathers, Griffin Smith, Zhitong Zhou
article en

Abstract

Abstract Background Generative AI video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions such as depression. Because AI-generated content may reach people during vulnerable periods, understanding what visual narratives these models produce for sensitive concepts carries clinical relevance. Objective This study aimed to characterize how OpenAI’s Sora 2 generative AI video model depicts depression and examine whether depictions differ between the consumer app and developer API access points, which differ in their product layer mediation. Methods We generated 100 videos using the single-word prompt “Depression” across 2 access points: the consumer app (n=50, 50%) and developer API (n=50, 50%). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Interrater reliability was assessed using the Cohen κ, with dimensions showing insufficient agreement excluded from analysis. Computational features (visual aesthetics, audio, semantic content, and temporal dynamics) were extracted and compared between modalities using 2-tailed Welch t tests with Benjamini-Hochberg false discovery rate correction. Results App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution compared with 14% (7/50) of API outputs. This divergence was reinforced across channels. App videos brightened over time (mean slope 2.90, SD 2.43 per second vs −0.18, SD 1.24 per second for the API; Cohen d =1.59; q <.001) and contained 3 times more motion (Cohen d =2.07; q <.001). Across both modalities, videos converged on a narrow visual vocabulary: predominantly seated figures (94/100, 94% of the videos); downward gaze (93/100, 93% of the videos); and recurring objects including hoodies (n=194), windows (n=148), and rain (n=83). Transcript language in depressive phases emphasized weight and containment (“heavy,” “drowning,” and “room”), whereas recovery phases reversed these patterns: brightness increased by 26.6% (Cohen d =0.68; P <.001), gaze shifted upward in 67% (30/45) of recovery videos, and terms such as “light” and “breath” emerged. Figures were predominantly young adults (323/367, 88% aged 20-30 years) and nearly always alone (360/367, 98%). Gender varied by access point: app outputs skewed male (34/50, 68%), and API outputs skewed female (33/56, 59%). Conclusions Sora 2 does not invent new visual grammars for depression but compresses and recombines cultural iconographies, whereas platform-level constraints substantially shape which narratives reach users. Clinicians should be aware that AI-generated mental health video content reflects training data and platform design rather than clinical knowledge and that patients may encounter such content during vulnerable periods.

Journal of Medical Internet ResearchVol. 28
Openalex Percentile: Top 9%
Digital Mental Health Interventions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.