DMA approximates factor-to-variable messages directly, bounds marginal KL by message KL, and yields a Bayesian neural network inference algorithm without a learning-rate hyperparameter
Synopsis
The work introduces Direct Message Approximation (DMA), which instead of approximating the marginal at each factor edge as expectation propagation (EP) and variational message passing (VMP) do, approximates factor-to-variable messages directly; for normalisable factors it defines a consistency condition (requiring exactness when all other incoming messages are Dirac deltas), proves a master theorem bounding marginal KL from message KL for proper messages on any graph with three structural corollaries—Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages—and additionally proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor; as a concrete instantiation it derives explicit DMA messages for the prod
Figure 2: Left : Posterior predictive mean (solid) and ± 2 σ \pm 2\sigma intervals (shaded) versus the true data-generating function (dashed), where σ = Var [ f ( x ) ] + β 2 \sigma{=}\sqrt{\mathrm{Var}[f(x)]+\beta^{2}} combines posterior variance with observation noise ( β = 0.2 \beta{=}0.2 ). Training points shown as crosses ( N = 200 N=200 , x ∈ [ − 2.5 , 1.5 ] x\in[-2.5,\,1.5] ); dashed verticals mark the training boundaries. Right : Hinton diagram of posterior weight beliefs after training. Square area ∝ \propto posterior mean magnitude; transparency ∝ \propto posterior variance (opaque = certain).
arXivInterpretation
Introduces Direct Message Approximation (DMA), which changes the object of approximation from the marginal at each factor edge to the factor-to-variable message itself, guided by a consistency condition: for normalisable factors, the result must be exact when all other incoming messages are Dirac deltas. EP and VMP both approximate the marginal, which forces an iterative round-robin schedule, risks negative-precision messages, and for VMP collapses to point estimates at Dirac-delta factors; DMA approximates messages directly and thereby avoids these structural sources at the level of construction. The paper defines the consistency condition and proves a master theorem (proper messages, any graph) bounding marginal KL from message KL, with three structural corollaries: Dirac-input consistency, no EP-style inner-loop iteration, and no negative-precision messages; the evidence is in the form of theorems and corollaries, and the abstract reports no experimental scale or numerical values.
Proves a complementary O(1/r^2) guarantee for the inherently improper backward message of the product factor. Closed-form treatment of this backward message has resisted prior work, and DMA provides a characterisation with an order guarantee. The abstract explicitly calls this a 'complementary O(1/r^2) guarantee' for the 'inherently improper backward message of the product factor' and notes that its closed-form treatment 'has resisted prior work'; specific constants and derivation details require the main text.
Derives explicit DMA messages for the product and leaky-ReLU factors and assembles a Bayesian neural network (BNN) inference algorithm with one forward/backward sweep per training example and no gradient learning-rate hyperparameter. Compared with approximate message passing routes that rely on iterative schedules and learning-rate tuning, this instantiation lands the structural guarantees in a concrete, runnable BNN inference procedure. The abstract gives the algorithm-level description (one forward/backward sweep, no learning-rate hyperparameter) and states that the structural guarantees were validated as predictive uncertainty that widens in data-sparse regions, including under model mismatch; no datasets, metric values, or comparison baselines are given.
Validates that the structural guarantees translate into predictive uncertainty that widens in data-sparse regions, including under model mismatch. This links the theorem-level KL bound to observable uncertainty behaviour in a BNN, rather than leaving it at formal derivation. The abstract states this validation as 'validating that the structural guarantees translate to predictive uncertainty that widens in data-sparse regions, including under model mismatch'; the concrete experimental setup and quantitative results are not given in the abstract.
Perspective
The work targets approximate inference on factor graphs, applies to normalisable factors on any graph, and uses the product and leaky-ReLU factors as concrete instantiations to assemble a Bayesian neural network inference algorithm with one forward/backward sweep per training example and no gradient learning-rate hyperparameter. It suits readers who want message-level approximation without EP-style inner-loop iteration and without negative-precision messages, and readers interested in how predictive uncertainty behaves in data-sparse regions and under model mismatch. The validation described in the abstract is qualitative, so the scope of applicability is bounded by the factors and the BNN setting the abstract explicitly covers.
The abstract gives no datasets, sample sizes, evaluation metrics, or numerical comparisons against baselines, so the strength of the validation that 'predictive uncertainty widens in data-sparse regions' can only be taken as stated in the abstract. The precise conditions, constants, and proof details of the master theorem and the O(1/r^2) guarantee require the main text. The explicit DMA message forms for factors beyond the product and leaky-ReLU, and how this BNN inference algorithm performs at larger scale or on different tasks, are directions a reader can continue to watch. From the abstract alone, it is not possible to judge how these guarantees behave in finite-precision implementations.
