Public articles linked to the same research event.
arXiv The work proposes DyRA, an input-adaptive method that dynamically approximates and corrects the output residual error introduced by structured weight approximations (such as low-rank factorizations) during inference by directly optimizing low-rank factors of the output, yielding a more faithful approximation of full matrix multiplication under the same computational budget; across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone, and notably achieves a 1.5x end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3x.
The work proposes DyRA, an input-adaptive method that dynamically approximates and corrects the output residual error introduced by structured weight approximations (such as low-rank factorizations) during inference by directly optimizing low-rank factors of the output, yielding a more faithful approximation of full matrix multiplication under the same computational budget; across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone, and notably achieves a 1.5x end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3x.
The work proposes DyRA, an input-adaptive method that dynamically approximates and corrects the output residual error introduced by structured weight approximations (such as low-rank factorizations) during inference by directly optimizing low-rank factors of the output, yielding a more faithful approximation of full matrix multiplication under the same computational budget; across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone, and notably achieves a 1.5x end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3x.
The work proposes DyRA, an input-adaptive method that dynamically approximates and corrects the output residual error introduced by structured weight approximations (such as low-rank factorizations) during inference by directly optimizing low-rank factors of the output, yielding a more faithful approximation of full matrix multiplication under the same computational budget; across vision, speech, and language models, DyRA consistently improves the accuracy-efficiency trade-off over structured weight approximations alone, and notably achieves a 1.5x end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than 3x.