GREQ two-phase hybrid framework delivers up to 40% throughput gain and 100% QoS satisfaction in distributed 6G resource management
Synopsis
The work introduces GREQ, a two-phase hybrid model-learning framework: phase one uses a channel-statistics-based greedy algorithm to deterministically satisfy user QoS with minimal subchannels, and phase two exploits the OFDM property that inter-cell interference on a subchannel comes only from the same subchannel to decompose the large-scale MARL problem into independent per-subchannel QMIX subproblems, avoiding exponential action-space growth; in simulations with 50MHz bandwidth and different numerologies, GREQ improves throughput by up to 40% over optimization and learning baselines, achieves 100% constraint satisfaction, and activates only 20–80% of available subchannels.
Fig. 1 : Block-diagram of the proposed hybrid learning framework. Each BS only observes local CSIs and long-term network information.
arXivInterpretation
A two-phase hybrid framework is proposed that offloads QoS constraint satisfaction from learning to a deterministic greedy allocation based on channel statistics, leaving phase two to solve only unconstrained throughput maximization. Unlike risk-aware learning that directly minimizes violation probability via chance constraints or CVaR, this design assigns the hardest task, hard-constraint satisfaction, to a model-based procedure, sparing the learning phase from complex reward shaping and Lagrange multiplier tuning. In small-scale training, GREQ maintains a consistent zero violation rate while episode reward rises rapidly, whereas MAPPO-Lagrange attains significantly lower rewards with persistently higher violation rates, and its reward declines as it reduces violations.
Exploiting the OFDM structure that inter-cell interference on a subchannel is caused only by transmissions on the same subchannel, the global MARL problem is decomposed into independent per-subchannel subproblems, each trained with a separate QMIX instance and one agent per subchannel per BS. Prior work runs QMIX over the full joint action space, causing computational infeasibility, or avoids action-space explosion by restricting to TDD, limited power steps, or at most one subchannel per user; this decomposition reduces the action space from exponential to linear without performance loss. The text notes that with 3 BSs, 20MHz bandwidth, 120kHz numerology, and 10 users each, the joint action space cardinality is about 4 billion; in the large-scale setting MAPPO-Lagrange is excluded due to the prohibitive action space.
Invalid action masking encodes phase-one assignments and idle serving-slot constraints directly into each agent's feasible action set, converting a value-based constraint into an action-set constraint. This avoids complex constraint-handling mechanisms and enforces constraints through a lightweight, widely used masking technique, improving training stability and implementation reliability. The design is described at the method level and works together with the phase-one greedy assignment and the fixed serving-slot abstraction, supporting the near-zero QoS violations observed in training and testing.
In 3GPP-aligned simulations, GREQ consistently outperforms optimization and learning baselines in throughput, QoS satisfaction, and resource utilization, and learns to selectively deactivate subchannels to mitigate inter-cell interference. Baselines RRS, RAND, and DOPT activate all subchannels regardless of load, yielding roughly 30% lower throughput and failing to meet every user's QoS; GREQ activates only 20–80% of subchannels and, since per-subchannel power is fixed, also achieves more efficient power usage. Both small-scale and large-scale settings report GREQ with the highest throughput, near-zero QoS violations, and least resource usage; in the large-scale setting GREQ activates only 20–70% of subchannels, while RRS and DOPT maintain violations around 3%.
Perspective
The results target multi-cell downlink OFDM systems with a bounded number of serving slots per BS, where users arrive and depart randomly and their locations and QoS fluctuate, under shared spectrum and no central coordination. Phase one relies on interference channel statistics (pathloss) and the assumption that all subchannels at other BSs are active, and adds a safety margin ratio to reduce the risk that the true rate falls short of QoS. For practical deployment, the authors state the trained model should be fine-tuned with real data from a measurement campaign and can be adjusted in real time via multi-modal or online learning in the Open RAN RIC module; for large networks sharing spectrum, a divide-and-conquer clustering strategy can be used, where each BS must also estimate interference statistics from other clusters.
All results are obtained in simulation, so model mismatch in real propagation environments and the effect of fine-tuning with measurement data remain to be verified. The choice of the safety margin ratio is currently based on numerical experience, and the authors explicitly state that mathematical analysis is not considered, so its robustness across different interference distributions is worth watching. In the large-scale setting, MAPPO-Lagrange is excluded due to the prohibitive action space, so direct comparison with constrained MARL at that scale remains an open question. The impact of cross-cluster interference statistics estimation accuracy on performance in clustered deployment also awaits investigation.
