ARGUS audits identification assumptions in climate-policy DID studies, detecting 8 of 11 injected flaws and abstaining on about 60% of paper-dimension assessments across 26 economics papers
Synopsis
The authors introduce ARGUS, a bounded retrieval-gated large-language-model pipeline that decomposes a difference-in-differences (DID) study into eleven assumption-implication-evidence dimensions, audits whether the evidence a paper reports adequately supports each identification assumption, and abstains when no relevant evidence can be retrieved; it detects 8 of 11 planted flaws (versus 2 for a keyword baseline), 25 of 33 flaw variants, abstains on roughly 60% of paper-dimension assessments over 26 economics papers for lack of retrievable evidence, and in a five-paper, 55-cell pilot with two reconciled annotators is more severe than the human labels on 25 of the 33 cells it completes, an over-severity a rule fixed before the labels arrived reduces substantially in-sample.
Interpretation
ARGUS operationalizes DID identification credibility as eleven auditable dimensions (parallel trends, no anticipation, staggered-timing handling, SUTVA and spillovers, control-group construction, specification, inference, sample-period selection, concurrent policies, placebo/robustness, data measurement), each written as an assumption, a testable implication, and the evidence a credible paper would report, stored declaratively so the audited dimensions are explicit and inspectable. The econometrics literature already formalizes when DID estimators fail under staggered timing and heterogeneous effects; what is added here is turning each known threat into a dimension whose supporting evidence can be checked in a given paper. The paper states the rubric maps one-to-one onto a flaw taxonomy, and each injected flaw is traced to the literature that names it (Appendix B, Table 3), so the threat catalogue is inherited while the injector's prose is the authors' own.
On the 11-flaw benchmark the two-stage ARGUS detects 8 flaws with zero false alarms against 2 for the keyword baseline; on the 33-variant benchmark the two-stage pipeline detects 25, a per-dimension no-retrieval ablation detects 29, and a single full-paper call detects 24. The authors replace sentinel-sentence injection with structural perturbation: omission flaws delete the section, figure, or table supplying a dimension's evidence, and commission flaws rewrite a section in prose resembling a flawed paper without the keyword detector's predefined negative-signal phrases, with a regression test guarding against phrase leakage. The paper reports that an initial implementation inflated detection through phrase leakage and was replaced; three two-stage runs on the 33 variants are near-deterministic (identical detection on 32/33, risk labels on 31/33), but the authors note no pairwise difference between architecture arms reaches significance at n=33.
Across 26 economics papers tagged DID, about 60% of the two-stage pipeline's paper-dimension assessments are unknown, meaning the relevance gate could not surface evidence and the system abstains; where retrieval is weak, judgements tend to be high risk. This identifies retrieval rather than judgement as the coverage bottleneck and separates a retrieval failure from a paper genuinely lacking evidence: a failed gate short-circuits to unknown without calling the judge. The paper reports the corpus is the 26 DID-tagged papers from the 259-paper CausalVerify release, and notes the tag was not verified paper by paper, one entry is a literature review, and none is a climate-policy evaluation in the narrow sense.
In a five-paper, 55-cell pilot with two reconciled annotators, ARGUS abstains on 22 cells and is more severe than the human labels on 25 of the 33 cells it answers, never less severe; a rule fixed before the labels arrived, acting on weak-retrieval high cells, raises exact agreement from 0.24 to 0.76 and cuts over-severe cells from 25 to 8. The authors localize the calibration failure as treating 'weak retrieval found nothing' as a substantive high, and supply a reproducible demotion rule, while stating the lift is in-sample because the rule was derived on the same papers. Two annotators outside the author team labelled independently and reconciled 13 differing cells themselves, agreeing exactly on 45 of the 54 cells both rated; the authors stress this is a pilot-scale signal, with the 55 cells nested in five papers and the weighted kappa interval including zero.
Perspective
The method is aimed at empirical researchers, journal reviewers, and policy analysts who need to check the identification evidence of a DID study, in settings where a paper reports evidence and the task is to localize its weakest dimensions; the authors state that climate-policy evaluation motivates the rubric, while the present evaluation uses synthetic environmental-policy fixtures and a general-economics corpus, making a dedicated climate-policy corpus run the natural next test. Output is study-dimension-risk triplets with evidence-linked rationales for a human to make the final call, not an adjudication of the causal effect.
Open questions the authors list include: the human gold is pilot-scale (five papers, 55 cells, two annotators reconciling their own differences), 23 of its cells judge an analogue of the dimension, and its labels are mostly medium, which makes exact agreement lenient; no pairwise difference between architecture arms reaches significance at n=33 and the human-gold intervals are wide; the calibration rules were derived on the same pilot they are scored on, so the out-of-sample gain is untested; the rubric, flaw taxonomy, and injector are co-designed, so synthetic detection rates are best read as internal consistency under known threats; the parallel-trends dimension is a reporting check that does not test whether a pre-trend test had power; and the real corpus contains no narrow climate-policy evaluation. In addition, several numbers in the loaded text appear blank at the abstract and table positions, so exact proportions should be read from the original tables and appendices.
