Public articles linked to the same research event.
arXiv This work develops a theoretical framework for the generalization behavior of Transformer-based neural quantum states under in-context learning (ICL): it proves that there exists a Transformer architecture whose pointwise prediction error (MSE) decreases inversely with both the number of in-context examples and the Transformer depth, and that the depth required to achieve this guarantee scales only linearly with system size (the number of particles in continuous systems or qudits in discrete systems); the authors further extend the analysis to full quantum states formulated as rank-one density operators, deriving MSE-based generalization bounds under physical constraints over both continuous and discrete domains, and corroborate the predicted scaling with numerical simulations.
This work develops a theoretical framework for the generalization behavior of Transformer-based neural quantum states under in-context learning (ICL): it proves that there exists a Transformer architecture whose pointwise prediction error (MSE) decreases inversely with both the number of in-context examples and the Transformer depth, and that the depth required to achieve this guarantee scales only linearly with system size (the number of particles in continuous systems or qudits in discrete systems); the authors further extend the analysis to full quantum states formulated as rank-one density operators, deriving MSE-based generalization bounds under physical constraints over both continuous and discrete domains, and corroborate the predicted scaling with numerical simulations.
This work develops a theoretical framework for the generalization behavior of Transformer-based neural quantum states under in-context learning (ICL): it proves that there exists a Transformer architecture whose pointwise prediction error (MSE) decreases inversely with both the number of in-context examples and the Transformer depth, and that the depth required to achieve this guarantee scales only linearly with system size (the number of particles in continuous systems or qudits in discrete systems); the authors further extend the analysis to full quantum states formulated as rank-one density operators, deriving MSE-based generalization bounds under physical constraints over both continuous and discrete domains, and corroborate the predicted scaling with numerical simulations.
This work develops a theoretical framework for the generalization behavior of Transformer-based neural quantum states under in-context learning (ICL): it proves that there exists a Transformer architecture whose pointwise prediction error (MSE) decreases inversely with both the number of in-context examples and the Transformer depth, and that the depth required to achieve this guarantee scales only linearly with system size (the number of particles in continuous systems or qudits in discrete systems); the authors further extend the analysis to full quantum states formulated as rank-one density operators, deriving MSE-based generalization bounds under physical constraints over both continuous and discrete domains, and corroborate the predicted scaling with numerical simulations.