Skip to main content
Back to timeline
arXivSource publication:

Machine Learning Models for Sequential Conditional Independence Testing: Error Decomposition Explains Why Exchangeability Tests Win in Practice

Synopsis

Addressing the difficulty of incorporating machine learning models into log-optimal e-variable design for conditional independence testing, this work decomposes the error into null enlargement, approximation, and estimation error, explains why exchangeability tests with lower theoretical power can have higher power in practice, explores intermediate null hypotheses between model-X conditional independence and exchangeability to reduce these errors, and provides estimation error bounds accommodating triple robustness results.

Source-provided article image: Sequential Conditional Independence Testing with Machine Learning Models
Figure 1 ·

Figure 1: While distilling information induces a null enlargement error , it can improve power by reducing the approximation and estimation errors .

arXiv

Interpretation

The paper proposes decomposing the error of conditional independence testing into null enlargement, approximation, and estimation error, using this to explain why exchangeability tests have lower power in theory but higher power in practice. Log-optimal e-variables had been studied, but how to incorporate machine learning models into their design remained unclear; this decomposition attributes the gap between theoretical and practical power to specific error sources rather than a general empirical observation. The evidence comes from an argumentative decomposition of error sources, a conceptual analysis framework; the source is a summary and provides no specific numerical experiments or sample sizes.

The decomposition shows that GRO e-variable estimates can be beaten because of their worse approximation and estimation errors. This offers a mechanistic explanation for why methods that directly test exchangeability outperform log-optimal e-variable approaches in practice, beyond merely reporting an empirical power difference. The conclusion is derived from the error decomposition; the source gives no scale or effect size for a specific comparison experiment.

The paper explores intermediate null hypotheses between model-X conditional independence and exchangeability to reduce approximation and estimation errors. Introducing intermediate null hypotheses between the two existing extremes opens a new design space for improving power while maintaining test validity. This is a method-design exploration; the source reports no construction details or empirical results for specific intermediate null hypotheses.

The paper notes that the model-X assumption often holds only up to an estimation error, invalidating exact type-I error guarantees, and provides estimation error bounds that accommodate triple robustness results with fast convergence rates. It explicitly incorporates the approximate validity of the model-X assumption into the error analysis and provides bounds accommodating triple robustness, offering a theoretical tool for error control when the assumption holds only approximately. The evidence is theoretical estimation error bounds and convergence rate results; the source gives no specific constants or numerical validation.

Perspective

The results target researchers working on conditional independence testing, e-variable design, and model-X methods, in settings that require trading off machine learning models against test power. The error decomposition framework and the exploration of intermediate null hypotheses offer directions for designing tests on concrete data, while the estimation error bounds provide a theoretical basis for controlling type-I error when the model-X assumption holds only approximately.

The source is a summary and does not give specific experimental setups, sample sizes, numerical results, or construction details of the intermediate null hypotheses, so the relative sizes of the decomposition components, the actual power gains from intermediate null hypotheses, and the specific constants in the estimation error bounds remain to be confirmed in the full text. The precise conditions for triple robustness and the exact meaning of the fast convergence rate also require consulting the full text.

Sources