How Generative AI Video Models Depict Depression: A Mixed Methods Study of OpenAI's Sora 2
Synopsis
Using the single-word prompt "Depression," this study generated 100 videos across two access points of Sora 2 (consumer app, n=50; developer API, n=50), had two trained coders independently code narrative structure, visual environments, objects, figure demographics, and figure states, and extracted computational features (visual aesthetics, audio, semantic content, temporal dynamics) for comparison, finding a pronounced recovery bias in app-generated videos (78%, 39/50, featured arcs progressing from depressive states toward resolution versus 14%, 7/50, of API outputs), with app videos brightening over time (mean slope 2.90, SD 2.43 per second vs -0.18, SD 1.24 per second for the API; Cohen d=1.59; q<.001) and containing three times more motion (Cohen d=2.07; q<.
Generalized system architecture for generative AI video platforms, illustrating the distinction between the developer API and consumer app access pathways. Both pathways share core safety infrastructure (center), including input and output moderation. Consumer apps additionally route prompts through an app layer (prompt processing, user experience [UX] constraints, default clip parameters, and policy-aware transformations) before generation and through a product surface layer (feed curation, distribution controls, feature thresholds, watermarking, and distribution controls) after generation. The developer API bypasses these layers, accessing only the shared safety infrastructure and base model. CSAM: child sexual abuse material.
PubMedInterpretation
The study characterizes the visual narratives Sora 2 produces for the sensitive concept of depression, finding that it does not invent new visual grammars but compresses and recombines existing cultural iconographies. Little was previously known about how generative video models represent conditions such as depression; this study fills that gap with coding and computational feature analysis of 100 videos. Two trained coders independently coded the material, with interrater reliability assessed using Cohen κ and dimensions showing insufficient agreement excluded from analysis; computational features were compared using 2-tailed Welch t tests with Benjamini-Hochberg false discovery rate correction.
Different access points to the same model produce markedly different narratives: the consumer app shows a pronounced recovery bias, while the developer API rarely does. The study identifies platform-level mediation (product layer differences) as a key variable shaping which narratives reach users, beyond the model itself. The 78% (39/50) versus 14% (7/50) narrative arc difference is reinforced across channels, including brightness slope (Cohen d=1.59; q<.001) and motion (Cohen d=2.07; q<.001).
Videos across both access points converge on a narrow visual vocabulary and figure setup. The study quantifies the degree of stereotyping in depression depictions, including seated figures (94/100), downward gaze (93/100), hoodies (n=194), windows (n=148), rain (n=83), and figures who are young (88% aged 20-30 years) and alone (98%). Based on systematic coding and object counts across 100 videos, with 50 videos per access point.
Depressive and recovery phases show measurable reversal patterns in language and visuals. The study maps narrative phases onto quantifiable changes in language, brightness, and gaze direction rather than describing them only qualitatively. Transcript language in depressive phases emphasized "heavy," "drowning," and "room," whereas recovery phases showed a 26.6% brightness increase (Cohen d=0.68; P<.001), upward gaze in 67% (30/45) of recovery videos, and terms such as "light" and "breath."
Perspective
The study delineates the depressive visual narratives Sora 2 generates under the single prompt "Depression" through two access points, the consumer app and the developer API, and is relevant to clinicians, platform designers, and researchers concerned with how AI-generated mental health content reaches users; its conclusions pertain to this model and these two access points rather than to all generative video systems.
Readers may still watch for: whether the single prompt "Depression" adequately represents broader depression-related expression; which coding dimensions were excluded for insufficient reliability; and whether similar recovery bias and visual vocabulary convergence appear in other generative video models or other access points.
