Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

GRAFT lets one backbone absorb five vision foundation models in sequence, with each new capability costing a single distillation

GRAFT introduces a continual multi-teacher knowledge distillation framework in which the previously distilled model acts as a teacher to preserve learned capabilities while the current student jointly learns from that model and the incoming teacher, and it adds Teacher Specific Readout Tokens plus a Geometry Agnostic Relational Loss to reconcile the incompatible representation geometries of heterogeneous teachers, yielding a single continually extensible backbone that unifies image understanding, 2D dense prediction, 3D human pose estimation, 3D vision, and vision-language, with each new capability acquired at the cost of a single distillation rather than a full re-distillation.