HyperNSDE writes static baselines into a neural SDE drift via a hypernetwork, improving observation-time fidelity on simulated and real clinical data
Related research and updatesSynopsis
The authors propose HyperNSDE, a continuous-time generative model that encodes heterogeneous static covariates with an HI-VAE, maps the static latent representation through a hypernetwork to subject-specific drift parameters driving a latent Neural SDE, and jointly models observation times with a latent-state-dependent intensity process, trained via a deterministic–stochastic path decomposition with a signature-kernel objective; on simulated data and the VELOUR and PPMI cohorts it achieves better observation-time fidelity and several distributional and privacy metrics, while forecasting and correlation performance is mixed and affected by observation-grid regularity and trajectory smoothness.
Figure 1: HyperNSDE generation pipeline. Generating a synthetic subject follows a hierarchical process (Appendix A.3 ). Static latent representations are sampled from the HI-VAE learned distributions, decoded into synthetic covariates, and passed through the hypernetwork to generate patient-specific SDE parameters. A latent path is then simulated from the personalized Neural SDE and decoded at observation times sampled from the intensity network to produce synthetic irregular longitudinal trajectories. Original longitudinal data are used only at training time via the loss terms, without relying on a longitudinal encoder. Solid arrows denote deterministic computations; dashed arrows denote stochastic steps.
arXivInterpretation
HyperNSDE places static covariates, irregular longitudinal trajectories, and observation times in one continuous-time generative framework, without a trajectory encoder or pre-imputation. The closest continuous-time joint model, MultiNODEs, injects static information mainly at the initial latent state and does not model observation intensity; here the static representation modulates the full temporal vector field through a hypernetwork, and the observation process is modeled explicitly. Table 1 compares representative models on static data, irregular time series, continuous latent dynamics, joint static/longitudinal modeling, and intensity modeling, with HyperNSDE the only one marked on all five; ablation AS.4 removes the observation model and reports a large increase in KL and MMD.
A deterministic–stochastic path decomposition with a signature-kernel objective stabilizes Neural SDE training on irregular sparse paths. Training a single latent Neural SDE to capture both smooth trends and stochastic fluctuations is reported as unstable; the latent path is split into a personalized Neural ODE trend term and a population-shared Neural SDE residual, trained with a reconstruction loss and a signature-kernel score respectively. Ablation AS.2 removes the decomposition and reports excessive stochastic noise overwhelming the trend signal, chaotic trajectories, poor color stratification, and a severe drop in correlation RMSE and utility; AS.1 replaces the SDE with an ODE and preserves utility but degrades correlation RMSE and MMD.
A latent-state-dependent intensity process lets generated data reproduce informative observation times, including an end-of-sequence event. Most joint generative models do not model observation intensity, or handle visit times only in discrete time; here neural networks derive per-variable observation intensities and a termination hazard from the latent path. On simulated data HyperNSDE achieves the best KL, discriminative, dependence, and predictive mean scores, with the lowest longitudinal MMD, lower than RTSGAN; on real data its KL is substantially lower than MultiNODEs, though RTSGAN obtains the lowest values.
Matched-grid analyses show that forecasting and correlation metrics are highly sensitive to observation-grid regularity and trajectory smoothness. Native-grid comparisons can favor smooth deterministic models; the authors reevaluate on common time grids and use a drift-only variant to isolate the smoothness effect. On the real observation grid, MultiNODEs' Corr. RMSE increases from one value to another while HyperNSDE changes from one value to another; on VELOUR prediction scores likewise shift with the grid, and HyperNSDE's drift-only prediction comes close to MultiNODEs.
Perspective
The work targets settings that need patient-level synthetic data: benchmarking predictive methods, studying rare subgroups, augmenting underrepresented cohorts, and sharing data under privacy constraints. The method applies to cohorts with heterogeneous static covariates, irregular and partially observed longitudinal variables, and informative observation times, and it learns the observational distribution of its training cohort. The authors state that in its present form it may support benchmarking use cases, while cohort augmentation or clinical trial simulation needs further progress; real-data experiments cover the VELOUR phase III metastatic colorectal cancer trial and the PPMI Parkinson's disease cohort, with code released and real data available through Project DataSphere and PPMI registration.
The evaluation protocol itself remains open: the authors note that no universally accepted protocol exists for irregularly sampled longitudinal data with missingness, that most benchmarks target regularly sampled or tabular data, and that interpolation-based preprocessing can distort temporal structure. Signature-kernel MMD depends critically on the bandwidth, and under sparse irregular sampling the interpolated paths are only rough proxies for the underlying continuous trajectories, so it should be interpreted with caution. On real data, forecasting and correlation metrics are affected by grid regularity and trajectory smoothness, longitudinal Corr. RMSE differences are not significant in paired tests, and paired KL results are baseline-dependent. The model learns the observational distribution of its training cohort rather than causal counterfactuals, so it cannot by itself correct selection bias or extrapolate to unsupported subgroups, and it is computationally more demanding than the baselines. The authors list richer diffusion parameterizations, typed longitudinal output distributions, differential privacy during training, and a deeper understanding of when the deterministic–stochastic decomposition helps as future directions.
