Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

FARM replaces policy transfer with reward-space transfer, reaching a 29.8% late-stage gain on unseen Rate-Latency tasks

The paper proposes FARM, a reward-space transfer framework that shifts cross-task knowledge reuse from policy space to trajectory-level decision evaluation: its Agentic Reward Model first learns a task-conditioned reward prior from heterogeneous source-task trajectories, then freezes it to guide adaptation of a target-specific controller for previously unseen tasks; on heterogeneous MEC tasks, FARM achieves a mean late-stage gain of 29.8% over Single-task SAC on unseen Rate-Latency targets, versus 16.1% for CRA Transfer, reaches a 46.2% gain on the moderate-OOD FAR-M case, and both Mamba and Transformer trajectory encoders support the transfer, with Mamba more robust as longer history dependencies are introduced.