Skip to main content
Back to timeline
arXivSource publication:

Cross-fitting still satisfies a central limit theorem under nonregularity, but its variance must be adjusted for cross-fold correlation

Related research and updates

Synopsis

For a common form of nonregularity in cross-fitting (testing whether a fitted model outperforms another, testing heterogeneous treatment effects with machine learning, and estimating the value of a potentially non-unique optimal treatment regime), the paper introduces a new locality condition and shows that a large class of cross-fitting estimators still satisfies a central limit theorem, but with an asymptotic variance that must be adjusted for cross-fold correlation; the author proposes a method to estimate this correlation and constructs new confidence intervals that attain asymptotically nominal coverage, with a simulation study using random forests and neural networks showing approximately nominal coverage.

AI-generated editorial illustration: Cross-Fitting Under Nonregularity: Normality and Inference via Locality

Interpretation

The paper states that conventional confidence intervals which ignore cross-fold dependence undercover in many applications sharing a common form of nonregularity, including the classic cross-validation problem of testing whether a fitted model outperforms another, testing for heterogeneous treatment effects with machine learning, and estimating the value of a potentially non-unique optimal treatment regime. Earlier discussions of cross-fitting validity largely concern regular settings; the paper groups these three applications under one form of nonregularity and identifies undercoverage of conventional intervals in that setting. This is a problem-scoping statement from the abstract, a theoretical positioning rather than a numerical or simulation result.

Using a new locality condition, the author shows that a large class of cross-fitting estimators still satisfies a central limit theorem despite the nonregularity, but with an asymptotic variance that must be adjusted for the cross-fold correlation. Where nonregularity might be expected to break normality, the paper shows a normal approximation remains available under the locality condition and identifies the adjustment as cross-fold correlation. This is the main theoretical result stated in the abstract, an asymptotic theorem; the abstract does not give proof details or the exact regularity conditions.

The paper proposes a method for estimating the cross-fold correlation and uses it to construct new confidence intervals that attain asymptotically nominal coverage. Building on the adjusted variance, it provides an operational variance-estimation and interval-construction procedure, turning the theoretical result into an inference recipe. The abstract presents the method as proposed by the author and claims asymptotic nominal coverage, a theoretical guarantee.

In a simulation study with random forests and neural networks, the proposed confidence intervals attain approximately nominal coverage. The theoretical intervals are examined on two widely used machine-learning estimators, adding finite-sample evidence beyond the asymptotic result. Evidence comes from a simulation study; the abstract reports no simulation size, coverage values, or data-generating settings.

Perspective

The result is aimed at researchers who use cross-fitting for inference, especially in economics and statistics applications that compare fitted models, test heterogeneous treatment effects, or estimate the value of an optimal treatment regime. It applies when the estimator satisfies the paper's locality condition and falls under the form of nonregularity described in the abstract; in that setting the proposed method estimates the cross-fold correlation and constructs confidence intervals that attain asymptotically nominal coverage, with simulations supporting approximately nominal coverage for random forests and neural networks.

The abstract does not give the exact locality condition, the form of the cross-fold correlation estimator, the additional regularity conditions needed for asymptotic nominal coverage, or the simulation sample sizes, data-generating processes, and coverage values. A careful reader would still watch whether the locality condition is easy to satisfy in real data and common machine-learning pipelines, how stable the correlation estimate is with few folds or limited samples, and whether the approximately nominal coverage in simulations extends to other estimators and more complex dependence structures. These are open questions rather than grounds for rejecting the conclusions.

Sources