Hasegawa and Ohzeki compare Ising and QUBO encodings under fixed model, sampler, and step size, finding QUBO has lower Fisher spectral entropy and slower SGD convergence while natural gradient descent makes the encodings converge alike
Synopsis
Under a controlled protocol that fixes the Boltzmann machine model, simulated-annealing sampler, and learning-rate design, the study compares Ising ({−1,+1}) and QUBO ({0,1}) variable encodings, exploits the identity that the Fisher information matrix equals the covariance of sufficient statistics to visualize empirical moments, and finds that QUBO induces larger cross terms between first- and second-order statistics, creating more small-eigenvalue directions and lowering spectral entropy, which explains slower convergence under stochastic gradient descent, whereas natural gradient descent, which rescales updates by the Fisher information matrix metric, achieves similar convergence across encodings due to reparameterization invariance.
Fig. 1.
· Page 5Interpretation
Under stochastic gradient descent, the Ising encoding consistently achieves faster Kullback–Leibler divergence reduction and requires fewer iterations to converge than QUBO on every dataset. Prior work on the Fujitsu Digital Annealer, QUBO transformation schemes for 3-SAT, and Ising-machine linear solvers had suggested that representation affects finite-time optimization behavior, but did not isolate the encoding effect under a fixed model, sampler, and step size for Boltzmann machine learning; this study fixes a fully connected Boltzmann machine, simulated-annealing sampler (β=1, 10,000 samples), and learning-rate design to attribute the difference to encoding itself. On BAS 2×2 (450 samples), BAS 3×3 (1,120 samples), and Ising-sampled synthetic data with d=10, Jc∈{0.5,1.0,1.5}, 2,000 samples each, the authors report that 'From Figures1-6, we did not observe QUBO converging faster than Ising on any dataset,' with ηSGD=0.01/λmax(Fθ).
Natural gradient descent makes the Kullback–Leibler divergence trajectories similar across encodings, showing that rescaling updates by the Fisher information matrix metric restores representation invariance. The authors characterize switching from QUBO to Ising as equivalent to centering and rescaling that acts as Fisher information matrix preconditioning, and test this information-geometric account under controlled conditions using a damped natural gradient step θt+1=θt+ηNGD(Fθ+λI)−1∇L(θt) with λ=0.001 and ηNGD=0.01. On the same datasets and sampling setup, the authors report that 'Under NGD, KL-divergence trajectories are similar across encodings' and that 'Proper curvature scaling (NGD) allows QUBO to achieve convergence comparable to Ising.'
QUBO's Fisher information matrix has systematically smaller eigenvalues and lower spectral entropy, originating from persistent non-zero cross-block correlations between first- and second-order sufficient statistics. Using the identity that the Fisher information matrix equals the covariance of sufficient statistics, the authors visualize the encoding difference as empirical moment distributions: Ising's first-order moments E[si] and third-order moments E[sisksl] are distributed around zero and its second- and fourth-order moments are symmetric around zero, making the cross block Fh,J very small and the Fisher information matrix approximately block-diagonal; QUBO's x∈{0,1} asymmetry yields non-negative moments with E[xi]≈0.5 and E[xixj]≈0.25, making the cross block non-zero. A block-form Fisher information matrix and the Schur bound λmin(Fθ)≤λmin(F22−F21F11−1F12) show that F12=F21≈0 for Ising while non-zero for QUBO, so the latter tends to produce smaller eigenvalues; Figure 3 shows Ising's spectral entropy is consistently larger, Figure 4 shows QUBO splits into large and very small eigenvalues early in training, and Figure 5 shows a smaller Jc prolongs the period with extremely small eigenvalues.
For QUBO variables, centering/scaling or natural-gradient-style preconditioning mitigates curvature pathologies, yielding actionable guidelines for variable encoding and preprocessing. The authors note that because a Boltzmann machine's Fisher information matrix equals the covariance of sufficient statistics, initialization and preprocessing that align first- and second-order moments can substantially reduce training iterations, especially for QUBO variables; they list block-diagonal or low-rank Fisher information matrices, Kronecker-factored curvature (K-FAC), and stochastic inverse estimators such as Hutchinson, Lanczos, Shampoo, and AdaHessian as promising routes to scalable natural gradient descent. The recommendation rests on the trajectory and spectral analyses of the controlled experiments above, plus the observation that natural gradient descent requires computing or approximating F−1, whose cost scales with parameter dimensionality; the authors frame it as practical guidance and future direction rather than a completed large-scale validation.
Perspective
The results apply to a controlled comparison under a fixed fully connected Boltzmann machine, simulated-annealing sampler (β=1, 10,000 samples), and prescribed learning-rate design, covering BAS 2×2, BAS 3×3, and Ising-sampled synthetic data with d=10 and Jc∈{0.5,1.0,1.5}. For practitioners training Boltzmann machines or restricted Boltzmann machines with stochastic gradient descent, this means preferring the Ising encoding under SGD, or adopting centering/scaling or natural-gradient-style preconditioning when retaining QUBO; for readers studying information geometry, it offers a framework for understanding encoding choice as Fisher information matrix preconditioning. The authors further note that scalable natural gradient descent can be pursued via block-diagonal or low-rank Fisher information matrices, K-FAC, and stochastic inverse estimators such as Hutchinson, Lanczos, Shampoo, and AdaHessian.
The authors list future directions including moment-matching initialization schemes, adaptive damping policies informed by spectral monitoring, hardware-aware studies using quantum and digital annealers (including online calibration of the effective inverse temperature for quantum annealing–based RBM training), and extensions to deep energy-based models where variable representation interacts with hierarchical structure. Readers may watch whether these directions reproduce the encoding differences at larger scale, in deeper models, and on real annealer hardware; the experiments here are small-scale fully connected Boltzmann machines, and the cost of computing or approximating F−1 for natural gradient descent scales with parameter dimensionality, so its scalability remains to be verified.
