Skip to main content
Back to timeline
arXivSource publication:

Structured Claim-Level Discourse Representations for Dense Health Narratives

Synopsis

This work represents dense health narratives as tuples linking atomic claims with thematic aspects, stance, and a six-dimensional pragmatic profile, builds a benchmark of 1,191 manually annotated claims from 60 videos across GLP-1 weight-loss medications, testosterone replacement therapy, collagen supplementation, and intermittent fasting, and evaluates large language models on structured discourse prediction under different discourse context settings.

Source-provided article image: Structured Claim-Level Discourse Representations for Dense Health Narratives
Figure 1 ·

Figure 1: An illustration of our formal task: decomposing high-density narratives into structured tuples. This segment demonstrates multi-aspect entanglement and shifting rhetorical profiles, moving from prescriptive medical policy (Row 2) to hedged physiological outcomes (Rows 3–4) and attributed personal narrative (Row 6).

arXiv · Page 2

Interpretation

It introduces a claim-level structured discourse representation that jointly models each atomic claim with its thematic aspect, stance, and a six-dimensional pragmatic discourse profile (logical relationship, verifiability and intent, evidence basis, certainty, temporality, claim focus). Prior approaches largely rely on coarse topic-level, sentiment-based, or stance-oriented labels; this work decomposes discourse into analyzable tuples and shows that two claims sharing the same aspect and stance can diverge across all six pragmatic axes. Supported by a qualitative comparison of two TRT-domain claims sharing the aspect 'Fertility & HPTA' and negative stance yet diverging across all six axes, plus a full tuple example.

It constructs the first benchmark for structured claim-level discourse analysis of dense health narratives, spanning four health domains with 1,191 manually annotated claims and an average density of 13.22 atomic claims per minute. Existing benchmarks typically study these dimensions in isolation or in comparatively constrained discourse settings; this benchmark covers thematic, stance, and multidimensional pragmatic structure under one annotation scheme. 60 videos, roughly 90 minutes of discourse, and about 400 person-hours of annotation and reconciliation; statement extraction pairwise F1 of 0.93, atomic decomposition agreement of 0.92, aspect kappa of 0.83, stance kappa of 0.80, and 6-axis kappa of 0.76.

Evaluation shows that LLMs perform strongly on thematic categorization and stance prediction but degrade substantially on high-dimensional pragmatic profiling, and that different discourse tasks benefit from different forms of contextual reasoning. Thematic tasks reached 79.23% aspect accuracy and 91.55% stance accuracy under batch windows with definitions and keywords plus lightweight chain-of-thought; pragmatic tasks reached 86.44% mean accuracy under a local plus or minus 2 window with heuristic constraints, while explicit chain-of-thought dropped to 73.12%. Ablations with Gemini 2.5 Flash across six context settings and multiple prompting strategies, plus a stability check on the TRT data across temperatures 0.0 to 0.8 with three trials each, totaling 12 runs.

Cross-model evaluation indicates that open-weight models remain competitive on aspect and stance classification but degrade more substantially on high-dimensional pragmatic profiling. Under identical prompting and context settings, Gemini 2.5 Flash, Llama 3.3, Qwen 3.5 35B, and Gpt-oss 20B were compared, with Qwen 3.5 35B consistently approaching proprietary-model performance on aspect and stance tasks. Four models evaluated across four domains under the identical B2-CoT pipeline, reporting accuracy and Cohen's kappa.

Perspective

The framework targets English-language, YouTube short-form, single-speaker narrative discourse across four health domains, using transcripts as the unit of analysis; the representation is positioned as extensible rather than a fixed schema, and the authors suggest it may apply to other dense narrative settings such as financial advice, political commentary, scientific debate, and legal discourse, while leaving full cross-domain validation to future work. For researchers interested in claim-level discourse analysis, health information quality assessment, or annotation scheme design, the benchmark and tuple representation offer a starting point for reproduction and extension under the same setting.

A careful reader may still watch: the best pragmatic profiling results rely on heuristic constraints manually derived from annotation reconciliation, whose transferability to new domains or label schemas has not been tested; the benchmark uses transcripts as the unit of analysis and abstracts away multimodal signals such as visual demonstrations, on-screen text, and creator affect; the framework focuses on single-speaker narrative discourse and does not explicitly model turn-taking or cross-speaker stance dynamics; and since the loaded text is a fast parse, some tables and appendix details appear in textual form, so readers wanting to verify specific numbers and prompt templates may wish to consult the original figures and tables.

Sources