Sequential resetting procedures deliver horizon-independent FDR control for LLM watermark detection and financial backtesting
Synopsis
The authors propose sequential resetting procedures built on e-values and test supermartingales that reset the capital to one after each rejection and continue testing, thereby reporting multiple intervals containing null violations; a refinement rule reports only observations after the within-block minimum of the capital. Under independent null e-values the FDR bound is α, while under dependence between null and non-null data the bound is α times a logarithmic factor with anytime validity, demonstrated on LLM watermark detection and ES backtesting.
Figure 1 : The texts are from the prologue of Efron’s Large-Scale Inference ( Efron, 2010 ) , where two text blocks are replaced by watermarked model output (in italics) using Tournament-watermarked OPT-1.3B continuations. Each underbracket, illustrated by “ token ﹈ \underbracket{\,\mathrm{token}\,} ”, represents one rejection block from our sequential resetting procedure. A report is false if its rejection block contains no watermarked position, and is true if there is at least one watermarked position. The above example reports 92.5% of the AI-generated text, and none of the reported blocks is false. See Section 5.2.5 for further details.
arXivInterpretation
Introduces sequential resetting procedures: conditionally valid e-values are multiplied into a test supermartingale, a rejection is declared when the capital first reaches the threshold α^{-1} and the observations since the previous reset are reported, after which the capital resets to one and testing continues, so multiple violation intervals can be reported on one data stream. Classical e-process sequential tests target a global null, stop at the first rejection, and do not indicate which data points caused it; this procedure turns a single rejection into repeatable interval reports while preserving the test supermartingale property. The method is given in Definitions 1 and 2 and equations (7) and (8); the authors note that resetting and refinement require only O(t) additional work beyond the underlying e-process, namely updating the e-process value and recording its running minimum and the time of the last minimum.
Introduces a refinement rule: within each rejection block, the last time the capital attains its within-block minimum is taken as the left endpoint, and only observations from that time through the rejection are reported, dropping less informative early data in long blocks. The ordinary procedure's reported intervals cover the entire testing horizon except the last segment, whereas the refined procedure reports a non-covering collection of intervals that locates violations more precisely; the authors prove both share almost the same provable FDR bounds and therefore recommend the refined version in most applications. Refinement does not change rejection times (right endpoints); left endpoints are chosen retrospectively at rejection. The authors state the refined FDR is pathwise at least that of the ordinary version, yet Example 3 shows a common tight bound for both, tight up to a small factor in the independent-null case.
Establishes two FDR bounds: under mutual independence of null e-values and independence from alternative e-values (Assumption IND) the FDR bound is α; when dependence between null and non-null data is allowed, null times may be adapted, and anytime validity is required, the uniform FDR bound is α times a logarithmic factor, with an explicit bound via geometric random variables. The authors state this is the first establishment of nontrivial FDR bounds for repeated sequential procedures based on test supermartingales; the proofs require a new weighting function linking the starting and ending capital of all-null intervals to cumulative FDR, and a refined version of Ville's inequality whose difficulty is that the block-minimum location is not a stopping time. Theorems 1 and 2 give the bounds; Example 1 shows the independent-case bound cannot be improved to α; Example 2 shows that either anytime validity or allowing dependence already costs a logarithmic factor of the same scale; Example 3 shows the bound is tight under a particular null-alternative arrangement; Example 4 shows the result requires test supermartingales rather than general e-processes.
Validates on LLM watermark detection and ES backtesting: for Gumbel-max and Tournament watermarks, resetting substantially raises the proportion of alternative positions covered relative to one-shot detection, and the refined version trades slightly lower power for much lower coverage FDP and higher IoU; for NASDAQ Composite ES backtesting, the resetting procedure yields multiple warning blocks around the 2008 financial crisis and the COVID-19 period and later. The authors state this is the first repeated online reporting of adaptively refined intervals containing LLM watermark evidence with nontrivial finite-sample FDR guarantees that remain valid under optional stopping in the general setting, and they extend existing one-shot sequential ES backtests to continue after rejection. Watermark experiments use facebook/opt-1.3b, 20 fixed prompts, 500 paths per temperature, 600 tokens per path, at temperatures 0.75 and 1; backtesting uses NASDAQ data from January 17, 1996 through December 31, 2025, with the test supermartingale containing 6,539 out-of-sample points from January 3, 2000 through December 31, 2025; all empirical uniform realized FDP values are below the theoretical bounds.
Perspective
The results apply where a conditionally valid e-value is available at each time step and reports are intervals rather than single points, as in LLM watermark detection and financial ES backtesting. Under the independent-null setting the threshold α directly yields an FDR bound of α; under dependence and anytime validity the threshold α yields a uniform FDR bound of α times a logarithmic factor, and the authors recommend the refined version in most applications unless narrowing reported intervals is not of interest. For watermark detection the method is not tied to a particular scheme, requiring only a token-level score with a known distribution when no watermark is present; for backtesting, resetting lets a regulator obtain further warnings after a first rejection rather than a permanent verdict based on the first crisis. The authors also note manufacturing process monitoring and anomaly detection as potential application areas.
The authors note that empirical FDRs are typically well below the theoretical bounds, a typical situation for e-value-based methods, and that there is room to improve power if additional assumptions are available. The Theorem 2 bound is expressed through geometric random variables and the explicit bound is asymptotically sharp as α tends to zero, but the sharpness construction relies on a particular data-dependent arrangement of null and alternative times, so behavior on real data still needs case-by-case assessment. Refinement shortens reported intervals, which may remove alternative observations and make a rejection easier to become false, so a true report need not consist entirely of alternative observations. Coverage FDP is a time-wise (token-wise in watermark settings) measure distinct from rejection-wise FDP, and coverage FDR has no theoretical FDR control. In backtesting, the empirical forecast's rapid warnings relate to its rolling historical window and small day-to-day variation, and the authors explicitly state this does not imply that specification is generally less conservative or more frequently rejected. In addition, this is a full-text reading, but specific numerical details in figures and appendices are not fully reproduced in the main text, so readers needing exact numbers should consult the original figures and the public code.
