RRM-GPT generates 5G NR scheduling grants field by field with an autoregressive model and transfers learned behavior to an unseen scenario
Synopsis
The authors propose RRM-GPT, an autoregressive framework that generates radio resource management decisions field by field as a language model generates text: an encoder maps heterogeneous network observations into common tokens and a decoder emits each decision conditioned on the network state and the fields already committed; in a 5G NR case study, a single model generates complete scheduling grants spanning user selection, timing, link adaptation, resource allocation, and control signaling, and transfers learned behavior to an unseen scenario without adaptation.
Fig. 1: Overview of the proposed autoregressive framework for RRM. Network state is encoded, and feasible, standard-compliant decisions are generated field by field. The formulation supports function-specific models with separate weights and, at its broadest scope, a shared model across functions. Models are pretrained on network logs and post-trained for a deployment objective.
arXivInterpretation
An autoregressive framework, RRM-GPT, that models RRM decisions as field-by-field sequences, pretrained on large volumes of network logs and then post-trained for deployment-specific objectives. Learning-based RRM models are typically specialized to one function and one deployment, so each new setting repeats the pipeline of assembling data, designing an objective, training, and validating; this framework unifies decision generation as next-field prediction so one model can be reused across deployments. A methodological and vision-level contribution argued from a structural observation (decisions are assembled from interdependent, standard-defined fields) and a two-stage training procedure; no experimental evidence is given for cross-function reuse.
A unified representation of network observations and RRM decisions: categorical values through a learned dictionary, continuous values by their position within learned intervals to retain magnitude, and the occupancy map in patches to preserve layout, with related features grouped and projected into a common embedding space to form input tokens. Text input has a predefined vocabulary, whereas RRM input is heterogeneous and varies with the use case; this representation lets scheduling, interference coordination, and handover share one generation mechanism, with output token values drawn from 3GPP standards and network configuration. Table I sets out the correspondence item by item for language models and for scheduling, interference coordination, and handover across input, token, vocabulary, complete sequence, rules, and end-of-sequence, as a conceptual argument.
In a 5G NR case study, a single model generates complete scheduling grants covering user selection, data-slot offset, MCS, PRB count, starting PRB, control-channel aggregation level, and starting CCE, and captures dependencies among grant fields. Prior autoregressive Transformer work on scheduling decisions used a dedicated model for each scheduling subtask; here one model generates the whole grant in a fixed field order, and the chain-rule factorization is exact, imposing no independence assumption among tokens. About 6.5 million training grants and 637,000 held-out grants from the same scenario; timing, aggregation level, and retransmission MCS are reproduced with relatively high rollout accuracy, and uplink starting-PRB accuracy rises from 48% to 99% under teacher forcing with MAE falling from 2.8 to 0.1 PRB indices, which the authors read as error propagation accounting for much of the rollout mismatch.
The model transfers without adaptation to a static-user, 167 m-radius scenario excluded entirely from training, improving most reported metrics, though transfer is not uniform. This provides an initial basis for reuse across deployments: initial-transmission MCS accuracy rises from 40% to 68% for downlink and from 47% to 91% for uplink, indicating the model has not memorized the training scenario, while uplink starting-PRB accuracy falls from 48% to 37%, showing transfer is directionally uneven. The transfer test uses 158,000 grants from a scenario separate from training; the authors attribute part of the gain to more stable channel conditions under static users and a smaller cell.
Perspective
The framework addresses vendors, operators, and researchers who want to reduce repeated engineering effort: within the function-specific scope, one scheduling model can be pretrained on network logs and adapted to different cell configurations, environments, and traffic conditions; the broader ambition is for load management, handover, interference coordination, and scheduling to share one model, invoked when each function's control opportunity arises, with generation order following inter-function dependencies and slower decisions not regenerated every slot. The case-study conclusions apply to the simulated 3.5 GHz, 40 MHz, 106-PRB urban macro cell with 100 KB FTP traffic, and to the static-user, 167 m-radius scenario excluded from training. On feasibility, values that would violate current constraints are masked before selection during decoding, with a lightweight state tracker recording selected field values and resource commitments; the authors state that a model trained on sufficient logs largely respects these rules on its own so masking rarely changes the outcome, but that even rare violations can disrupt network operation, making such lightweight guarantees essential for practical deployment.
Cross-function model sharing is not experimentally validated; its value is argued from complementarity and substitutability among functions rather than measured. Post-training, both behavioral cloning and reinforcement learning, is described as the route to deployment adaptation, but the case study reports no post-training results and no effect on network KPIs such as throughput or latency. The transfer evaluation covers a single static small-cell scenario, and the authors state explicitly that broader generalization, the benefits of post-training, and the impact on network performance remain open for future evaluation. Which fields to use, their granularity, and their order remain design questions, and whether a shared representation can transfer across standard releases and across systems such as Wi-Fi or satellite is unresolved. Log coverage and detail, how to combine logs from different vendors and deployments, complexity and latency under per-slot inference budgets, consistency among model instances deployed at multiple points, and how to express operator intent and adapt to it without unsafe exploration in a live network are all listed as open problems. In addition, equation (1) and some symbols are not fully rendered in the provided text, so reproducing the chain-rule factorization details would require the original figures and equations.
