Skip to main content
Back to timeline
medRxivSource publication:

The Online Health Safety Gap: Consensus Alignment Does Not Imply Safety in Peer-to-Peer Health Narratives

Synopsis

This work identifies and empirically validates the Online Health Safety Gap: across 704 health narratives (699 classified), two scores computed from non-overlapping feature sets, Narrative Truth Distance for epistemic divergence and Narrative Risk Score for health risk potential, share only 4.9% of their variance (r=0.222, p<0.001), so 39.6% of narratives fall in the two off-diagonal quadrants that single-axis systems mishandle by construction (aligned-but-risky 25.2%, divergent-but-safe 14.4%), and on an expert-labeled misinformation benchmark of 437 posts containing 127 misinformation instances the highest discrimination is Youden J of 0.349, motivating a Classification Quadrant and a shift from fact-centric evaluation to risk-aware assessment.

Source-provided article image: The Online Health Safety Gap: Consensus Alignment Does Not Imply Safety in Peer-to-Peer Health Narratives.
Figure 1 ·

actionable. Figure 1 presents the partition.

medRxiv · Page 8

Interpretation

It identifies and empirically validates the Online Health Safety Gap, a structural misspecification in which medically accurate content can carry substantial health risk. Prior online health information systems treat alignment with medical consensus as a reliable signal of safety; this work shows that assumption fails in peer-to-peer health discourse and quantifies the failure. Based on 704 health narratives (699 classified), where two scores from non-overlapping feature sets share only 4.9% of variance (r=0.222, p<0.001), a correlational result.

It quantifies the scale of what single-axis systems mishandle: 39.6% of narratives fall in the two off-diagonal quadrants, with aligned-but-risky at 25.2% and divergent-but-safe at 14.4%. It turns the claim that alignment is not safety into a countable distribution, showing fact-checking misses risky content and incorrectly suppresses safe content. Quadrant distribution statistics from the same 704-narrative corpus, a descriptive counting result.

Existing fact-accuracy benchmarks cannot evaluate risk-aware assessment: on an expert-labeled misinformation benchmark of 437 posts containing 127 misinformation instances, the highest discrimination is Youden J of 0.349, achieved by a supervised classifier trained directly on those labels, which neither frozen biomedical embeddings nor a prompted large language model improves upon. It shows labels encoding factual accuracy carry little signal about behavioral risk, so benchmarks of this construction cannot evaluate risk-aware assessment. A discrimination comparison on an expert-labeled benchmark, contrasting a supervised classifier, frozen embeddings, and a prompted large language model.

It introduces the Classification Quadrant, mapping epistemic divergence and health risk potential onto four governance categories with distinct intervention implications, and argues for a shift from fact-centric evaluation to risk-aware assessment. It offers a governance framework to operationalize the gap, extending the evaluation target from a single factual dimension to a risk dimension. A conceptual framework and argumentative contribution, grounded empirically in the corpus and benchmark analyses above.

Perspective

The work addresses peer-to-peer health discourse and applies to online health information and content-governance systems that use alignment with medical consensus as a safety proxy, as well as to evaluation practices that build misinformation benchmarks; its empirical results rest on 704 health narratives (699 classified) and a 437-post expert-labeled benchmark containing 127 misinformation instances, and depend on the VERITAS implementation and biomedical knowledge bases such as UMLS, SNOMED CT, SemMedDB, and PubMed.

A careful reader would still watch how stable the independence of the NTD and NRS scores is across health topics, platforms, and languages; how the four Classification Quadrant categories are bounded and operationalized; and whether discrimination levels such as Youden J of 0.349 hold under other benchmark constructions. This was an incomplete reading, so figures and full method details are missing, and these questions need confirmation against the original text and companion prior work.

Sources