FARM replaces policy transfer with reward-space transfer, reaching a 29.8% late-stage gain on unseen Rate-Latency tasks
Related research and updatesSynopsis
The paper proposes FARM, a reward-space transfer framework that shifts cross-task knowledge reuse from policy space to trajectory-level decision evaluation: its Agentic Reward Model first learns a task-conditioned reward prior from heterogeneous source-task trajectories, then freezes it to guide adaptation of a target-specific controller for previously unseen tasks; on heterogeneous MEC tasks, FARM achieves a mean late-stage gain of 29.8% over Single-task SAC on unseen Rate-Latency targets, versus 16.1% for CRA Transfer, reaches a 46.2% gain on the moderate-OOD FAR-M case, and both Mamba and Transformer trajectory encoders support the transfer, with Mamba more robust as longer history dependencies are introduced.
Fig. 1: Overview of heterogeneous wireless tasks considered by FARM. The task space spans diverse channel conditions, QoS objectives, constraints, and deployment scenarios. (I) Rayleigh fading models dense urban and mobile MEC environments with rapid non-line-of-sight variations. (II) Rician fading represents line-of-sight-dominant UAV-assisted and air-ground MEC scenarios. (III) Nakagami- m m fading captures adjustable propagation severity in V2X and industrial robotic environments. These heterogeneous tasks support trajectory-level reward learning for Reward-Space Transfer.
arXivInterpretation
FARM shifts cross-task knowledge reuse from policy space to trajectory-level decision evaluation, carrying transferable knowledge in a task-conditioned reward prior rather than in a shared policy. Conventional multi-task and transfer reinforcement learning methods primarily share or transfer policies, coupling transferable knowledge with task-dependent action mappings; FARM instead transfers in reward space, decoupling the transferable part from task-dependent action mappings. The claim is supported by the framework design itself: the ARM learns a task-conditioned reward prior from heterogeneous source-task trajectories and provides auxiliary guidance for task-specific policy optimization.
FARM uses a two-stage procedure: in Stage I the ARM jointly models task conditions, temporal trajectory dependencies, and objective-dependent reward structures while each source task retains its own controller; in Stage II the learned reward prior is frozen and reused to guide adaptation of a target-specific controller for previously unseen tasks, without transferring source-task policies. The two-stage design separates learning the reward prior from adapting the target controller, so transfer occurs through the frozen reward prior rather than through source-task policies. The procedure is explicitly described in the abstract as the division of labor between Stage I and Stage II, at the method level.
On heterogeneous MEC tasks, FARM achieves a mean late-stage gain of 29.8% over Single-task SAC on unseen Rate-Latency targets, compared with 16.1% for CRA Transfer, and reaches a 46.2% gain on the moderate-OOD FAR-M case. These numbers provide a quantitative comparison of reward-space transfer against a single-task baseline and an existing transfer method, with FARM's late-stage gain exceeding that of CRA Transfer. The evidence comes from the heterogeneous MEC task experiments reported in the abstract, including comparison values against Single-task SAC and CRA Transfer.
Analysis shows that both Mamba and Transformer trajectory encoders support Reward-Space Transfer, while Mamba provides improved robustness as longer history dependencies are introduced. This indicates reward-space transfer does not depend on a single encoder architecture, and that encoder choice affects robustness under longer history dependencies. The evidence comes from the further analysis described in the abstract, comparing the two trajectory encoders under reward-space transfer.
Perspective
The work targets learning agents that must adapt across heterogeneous channel conditions, traffic patterns, QoS requirements, objectives, and operational constraints, especially multi-task wireless network optimization settings such as multi-access edge computing. Its setting is that each source task retains its own controller, the ARM learns a task-conditioned reward prior from source-task trajectories, and that prior is then frozen to guide adaptation of a target-specific controller for previously unseen tasks without transferring source-task policies. It is therefore relevant to readers who want to reuse decision knowledge without also transferring task-dependent action mappings, for example those working on multi-task and transfer reinforcement learning, wireless resource management, and MEC optimization. The abstract-level results concern unseen Rate-Latency targets and the moderate-OOD FAR-M case, indicating the framework is usable in those settings.
The visible text is the abstract and does not include experimental setup, number of tasks, training budget, baseline tuning, or statistical significance, so the exact evaluation conditions behind the 29.8%, 16.1%, and 46.2% gains still need confirmation in the full text. Whether the reward prior remains effective when source and target tasks differ more, and how the trade-off between the frozen prior and target-controller adaptation changes with the number of tasks, are open questions worth watching. The observation that Mamba is more robust under longer history dependencies also awaits the full text for its scope and the dependency-length threshold.
