Skip to main content
Back to timeline
Climatic ChangeSource publication:

Generative Debunking of Climate Misinformation: Automatically Producing 'Truth Sandwich' Rebuttals with Large Language Models

Synopsis

This study introduces a 'generative debunking' framework that integrates climate contrarian claim classification (CARDS) and fallacy detection (FLICC) into an LLM prompting pipeline, so that a model takes a climate myth as input and produces a debunking that follows the fact-myth-fallacy-fact ('truth sandwich') structure; the authors pair three prompting strategies of increasing complexity with GPT-4, Palm2, and Mixtral, and four authors (including one climate misinformation expert) rate 60 debunkings for 20 myths on fact, fallacy, and structure, finding that GPT-4 with a simple prompt and Mixtral with a structured prompt perform relatively well, that fallacy explanations score higher than facts, and that non-expert annotators show poor agreement with the expert on fact quality.

Source-provided article image: Generative debunking of climate misinformation
Figure 2

Figure 2: Overview of our dynamic prompting approaches. Left: Single prompt with dynamic fallacy prediction (FLICC) and example retrieval (CARDS). Right: Structured prompt with additional ReAct component (Fact 1) and FEVER evidence retrieval (Fact 2). External resources are shown as diamonds, and shared components between the two approaches are highlighted in green.

· Page 4

Interpretation

It is the first integration of fallacy detection into an end-to-end, psychologically grounded, structured debunking system whose output strictly follows the fact-myth-fallacy-fact truth-sandwich format. Prior automatic fact-checking largely reduced the problem to veracity classification or extracted explanatory facts from supporting/refuting documents; this work embeds CARDS contrarian-claim classification and FLICC fallacy detection into the prompting framework to generate free-text debunkings directly. The authors state that 'to the best of our knowledge we are the first to integrate fallacy detection into an end-to-end system for psychologically grounded, structured debunking,' and provide full prompt templates (Appendix Tables 5-11) and a system architecture diagram (Figure 2).

Across 20 unseen climate myths, all three model-prompt combinations followed the truth-sandwich structure 100% of the time, but fact quality was the main shortcoming. Structural compliance was not the hard part; generating facts that are both correct and directly relevant to the myth was. The authors identify 'a lack of factuality and relevancy' as a critical shortcoming even with the latest LLMs. Four authors (including one climate misinformation expert) independently and blindly rated 60 debunkings on a 1-3 scale for Fact 1, Fallacy, and Fact 2; the authors report 100% structural compliance for all models and therefore disregard that score going forward.

GPT-4 with a simple generic prompt and Mixtral with a structured context-sensitive prompt generally outperform Palm2 with a single context-sensitive prompt; Mixtral scores highest on fallacy explanation, while GPT-4 tends to perform better on fact generation, particularly as judged by the expert. This suggests structured prompting can bring an open-source model close to a proprietary LLM on debunking, offering a path toward low-cost, widely deployable automatic debunking. Table 3 reports averaged ratings from all annotators, non-experts (n=3), and the expert (n=1): GPT-4 fact average 2.28 and fallacy 2.44; Palm2 fact 1.91 and fallacy 2.20; Mixtral fact 2.25 and fallacy 2.55.

Non-expert annotators show poor agreement with the expert and are systematically more optimistic when evaluating debunking facts, and they are easily persuaded by specific but not necessarily relevant facts. This reveals that evaluating automatic debunking systems itself requires domain-expert involvement or close supervision, because non-expert users are unlikely to detect the system's flaws. Table 1 shows generally low agreement on facts (e.g., Mixtral Fact 1 with expert Gwet AC1 = 0.09) and substantially better agreement on fallacies (GPT-4 fallacy with expert Gwet AC1 = 0.66); the authors give examples where non-experts gave high scores to Mixtral facts containing specific details such as 'only 0.3°C' while the expert judged them irrelevant.

Perspective

The framework targets the research setting of automatic debunking of climate myths and is suited to exploratory use where rebuttal texts need to be generated at scale in the fact-myth-fallacy-fact structure; the authors explicitly position it as a research tool for collecting user experiences in a controlled environment and stimulating follow-up work, not for broader deployment. For readers who want to reuse the approach, its value lies in the combination of prompt design, external knowledge retrieval (FLICC, CARDS, CLIMATE-FEVER), and the evaluation rubric, rather than in directly producing publishable content.

The authors do not systematically disentangle the contributions of model choice versus prompting strategy, nor do they exhaustively combine all prompts with all LLMs, so it is hard to tell whether performance gains come from model capability or prompt design; the evaluation covers only 20 myths and four annotators, and it does not systematically examine how models perform on common versus novel myths, though the authors observe that stronger models do well on widespread myths and sometimes generate irrelevant facts or misclassify fallacies for less common ones; moreover, the study assumes all input is non-factual and does not evaluate the models' ability to distinguish myths from facts. Readers may watch how follow-up work improves the specificity and relevance of facts and how expert supervision can compensate for the limits of non-expert evaluation.

Sources