Public articles linked to the same research event.
arXiv Addressing the unclear behavior of modern selective state space models such as Mamba2 in distributed learning, the work derives architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization; it then numerically validates the single-layer bounds on sequences generated by a teacher SSM, uses the analysis to formulate expectations about local training and client heterogeneity, and examines these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains.
Addressing the unclear behavior of modern selective state space models such as Mamba2 in distributed learning, the work derives architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization; it then numerically validates the single-layer bounds on sequences generated by a teacher SSM, uses the analysis to formulate expectations about local training and client heterogeneity, and examines these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains.
Addressing the unclear behavior of modern selective state space models such as Mamba2 in distributed learning, the work derives architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization; it then numerically validates the single-layer bounds on sequences generated by a teacher SSM, uses the analysis to formulate expectations about local training and client heterogeneity, and examines these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains.
Addressing the unclear behavior of modern selective state space models such as Mamba2 in distributed learning, the work derives architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization; it then numerically validates the single-layer bounds on sequences generated by a teacher SSM, uses the analysis to formulate expectations about local training and client heterogeneity, and examines these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains.