Skip to main content
Back to timeline
arXivSource publication:

Without knowing the covariate distribution, a new stochastic-approximation estimator attains parametric rates in logistic regression with missing covariates and provably beats complete-case analysis

Related research and updates

Synopsis

For logistic regression with an unknown (but bounded) covariate distribution and coordinate-wise missing-completely-at-random covariates, the authors construct a distribution-free estimating equation from a monotone operator and solve it with a projected stochastic approximation algorithm; they prove parametric-rate signal recovery with risk characterized by the missingness profile, show the estimator is never worse than the complete-case estimator, and provide a matching minimax lower bound under homogeneous missingness showing the nonstandard dependence on the observation probability is intrinsic.

Source-provided article image: Assumption-lean logistic regression with missing covariates
Figure 1 ·

As an illustration of the difficulties in this problem, Figure 1 considers a well-specified logistic model with bounded, non-Gaussian covariates in dimension d and independent coordinate-wise homogeneous MCAR missingness in which each coordinate is observed i.i.d. (and independently of the data) with probability q. We consider two settings, viz. (d, q) = (10, 0.40) and (100, 0.85); full details of this experiment are provided in Section C.

arXiv · Page 4

Interpretation

Proposes an estimator that neither models nor imputes the covariate distribution: it builds a population monotone operator whose unique zero is the true regression parameter, constructs an unbiased estimate of that operator from each incomplete observation, and embeds it in a projected stochastic approximation algorithm. Existing approaches rely on a known or estimable covariate distribution (imputation, EM, Gaussian working models) and can be inconsistent in nonlinear logistic regression even with large samples; this work formulates a distribution-free estimating equation and exploits the multiplicative factorization of the exponential term across coordinates to build the unbiased estimate. The paper establishes strong monotonicity and the zero property of the population operator (Lemma 5), conditional unbiasedness and variance control of the empirical operator (Lemma 6), and derives a finite-sample upper bound (Theorem 1); the simulation in Figure 1 shows imputation and Gaussian-likelihood procedures are inconsistent with bounded non-Gaussian covariates while the proposed algorithm approaches the complete-data MLE.

Provides a non-asymptotic risk upper bound whose missingness penalty is characterized by the observation-probability vector through an optimization problem, and shows the estimator is never worse than the complete-case estimator in any regime. Complete-case analysis retains only fully observed samples, with a missingness penalty of roughly the inverse product of coordinate observation probabilities; the bound shows the proposed method's penalty is never higher and can be substantially lower in high-dimensional homogeneous-missingness settings. Theorem 1 gives a parametric-rate bound (scaling as 1/n in the sample size) under coordinate-wise MCAR, bounded covariates, and a curvature condition; Corollaries 2 and 3 organize the homogeneous-missingness bound into three regimes depending on the relative magnitudes of dimension, curvature, and observation probability.

Establishes a minimax lower bound that matches the upper bound up to constants in the exponent under homogeneous missingness and restricted curvature, showing the nonstandard dependence on the observation probability is fundamental. In linear regression the missingness penalty is typically of order the inverse product of observation probabilities, whereas the logistic upper bound exhibits three distinct regimes; the lower bound shows this more intricate dependence cannot be removed by a better estimator. Theorem 4 gives the lower bound under a sample-size condition (Eq. 25), proved by combining two Assouad constructions: one capturing the intrinsic difficulty of the fully observed problem, and one using Fourier coefficients of a Boolean function (the majority function) to design a coupling that makes the observed-data likelihood ratio exactly 1 whenever too few auxiliary covariates are observed, thereby hiding the parameter perturbation.

Perspective

The results apply to a missing-completely-at-random mechanism in which the response is always observed, covariates are missing independently across coordinates with known observation probabilities, covariates are bounded, the model is a well-specified logistic regression, and the parameter lies in a bounded set. In this setting the method is directly usable in online or streaming scenarios: each iteration needs only one incomplete sample, and no covariate distribution needs to be estimated. For practitioners this means that when complete cases are scarce (for example, high dimension or low observation probability), incomplete samples need not be discarded while a parametric rate is retained; the explicit constants and three-regime characterization under homogeneous missingness also help judge when the gain over complete-case analysis is substantial.

The matching of upper and lower bounds holds under homogeneous missingness with restricted curvature; under heterogeneous missingness only the upper bound is available. The lower bound carries a sample-size condition, which the authors explain corresponds exactly to the regime where the upper bound is non-trivial. The method assumes known observation probabilities, bounded covariates, and correct model specification, and does not address sparse or other low-dimensional structure, misspecified models, MNAR, or more general nonlinear regression, all of which are listed as open problems. The simulation covers a single setting with bounded non-Gaussian covariates and two missingness probabilities, so practical performance still needs testing across more data-generating mechanisms.

Sources