Skip to main content
Back to timeline
arXivSource publication:

ALICE estimates mutual information zero-shot with one pretrained Transformer, cutting error at least twofold at the 1k-sample budget on the Beyond Normal benchmark and reproducing prior findings in biology, genetics, and neuroscience data

Synopsis

The authors present ALICE, a foundation model trained once on a synthetic distribution family that turns mutual information estimation into in-context estimation of rectified-flow velocity fields for unseen distributions, matching per-distribution neural estimators on the Beyond Normal benchmark with at least twofold lower error at a 1k-sample budget and zero-shot reproducing and extending dedicated estimators' findings on biology, genetics, and neuroscience data the model never saw.

AI-generated editorial illustration: ALICE: In-context, Zero-shot, Mutual Information Estimation

Interpretation

ALICE recasts mutual information estimation as a time integral of rectified-flow velocity differences and uses one context-conditioned Transformer to produce the joint field and the two block-conditional fields at once, removing per-distribution training. Earlier estimators such as MINE, InfoNCE, NWJ, SMILE, and the diffusion-based MINDE, InfoBridge, and FMMI must be retrained for every distribution, while InfoAtlas amortizes with a hypernetwork at a non-negligible accuracy penalty; ALICE keeps a single network conditioned on samples through attention and produces no distribution-specific parameters. The paper derives the velocity-form identity (Theorems 1 and 2) and states it is exact for the true fields and involves no variational bound; the training objective is a masked flow-matching MSE that requires no mutual information labels.

On the Beyond Normal benchmark, ALICE Base attains the lowest mean absolute error at all three matched budgets of 1k, 10k, and 100k samples, and at 1k samples its error is at least twofold lower than every competitor. Competitors on this benchmark are trained and tuned separately on each distribution while ALICE is zero-shot; the paper reports that at 1k samples the best classic estimator CCA reaches 0.42 nats and MINDE only falls from 1.51 to 0.40 nats between 1k and 100k, marking the small-sample regime as the weak point of existing neural estimators. Evaluation uses the Czyż et al. 2023 Beyond Normal suite with joint widths from 2 to 50 and closed-form ground truth; every estimator receives the same sample budget, ALICE splits it into query samples and a context of the remaining clean samples, and results are averaged over eight independent context draws.

In three scientific domains the model never saw, ALICE zero-shot reproduces dedicated estimators' conclusions and adds finer decomposition: single-cell NF-κB signaling, Arabidopsis TATA-box promoter localization, and O-information across six mouse visual-cortex areas. Jetka et al. 2019 approximated mutual information by fitting a linear classifier per analysis, and Bounoua et al. 2024 trained a score-based estimator on all sessions pooled; ALICE uses a frozen checkpoint to estimate each session separately, separating the familiar-image and novel-image days that the pooled analysis cannot distinguish. All three applications sit in the small-data regime: 100 to 500 cells per dose, roughly 10,000 sequences in the promoter dataset, and about 100 to 200 independent flashes per session; the paper reports capacity peak positions consistent with the reference study, a TATA-box peak inside the documented −40 to −20 base band, and positive O-information in every window and every mouse in novel-image sessions.

ALICE's architecture keeps parameter count independent of joint width and context length and natively supports different data dimensionality and sample cardinality, so one checkpoint serves discrete sequences, time series, and continuous vectors. InfoAtlas's hypernetwork emits fixed-shape weights that bind the estimator to the joint widths it was trained on; ALICE uses coordinate-shared scalar tokenization, set attention without positional encodings, an induced-latent bottleneck, and a context-derived relation graph, making parameter count independent of width and context length. The paper details the architecture and caching mechanism, noting that the relation graph, context tokens, and induced latents depend only on the context and are reused across queries; the promoter application encodes each base as one real coordinate and handles block dimensions of 4 and 600 natively.

Perspective

This work targets settings where per-distribution training is impractical and sample sizes are only a few hundred to a few thousand pairs, such as single-cell signaling, regulatory sequences, and neural recordings in biology and neuroscience. It lets researchers estimate mutual information per session, per dose, or per window with one frozen checkpoint instead of training a network for each system; the paper reports ALICE can run as a local model on modest hardware with a few forward passes per estimate. The applicability assumes data can be represented as coordinate blocks and that discrete symbols are mapped to real coordinates through a fixed injective embedding.

The limitations section states the implementation is research code that has not been thoroughly optimized, and notes that model size, training budget, training-corpus dimensionality, and the maximum training context length can all be increased. Two systematic biases remain on the benchmark: spiral embeddings are underestimated and dense multinormal tasks are overestimated, both growing with width, and ALICE Small degrades on wide tasks as the context grows. The three scientific applications have no ground-truth mutual information and their conclusions are drawn against the reference studies' biological findings; the neuroscience part also shows the familiar-image and novel-image sessions already differ by about 0.1 nats before the visual response arrives, and animal identity explains about 60% of the variance of peak-window estimates. In addition, the loaded text is missing many numeric entries in the benchmark result tables, so specific error values could not be verified and only the prose statements were used.

Sources