Skip to main content
Back to timeline
arXivSource publication:

Mirror-IRS-assisted hybrid RF/VLC network: two-agent DRL jointly tunes selection, power and mirror orientation, holding fairness above 0.93

Synopsis

For an indoor IRS-assisted hybrid RF/VLC downlink network, this work formulates joint proportional-fairness maximization over RF/VLC technology selection, power allocation and IRS mirror roll/yaw orientation as an MDP and solves it with a cooperative two-agent DRL framework under centralized training with decentralized execution, where one agent handles selection and power and the other handles mirror orientation; simulations converge to 350 Mbps, within 10% of the 392 Mbps exhaustive-search bound, outperform DDPG and DQN, and keep Jain's fairness index above 0.93 as user count grows.

Source-provided article image: Network Adaptation in IRS-Aided Hybrid RF/VLC Systems Using Cooperative Multi-Agent DRL
Fig. 1 ·

Fig. 1: IRS-assisted RF/VLC system model with DRL framework.

arXiv

Interpretation

Formulates RF/VLC technology selection, power allocation across both subnetworks, and IRS mirror roll/yaw orientation as a single proportional-fairness maximization problem, and shows it is a non-convex mixed-integer nonlinear program. Prior hybrid RF/VLC work often treats the VLC channel as fixed and uncontrollable, or configures the IRS with model-based methods such as a sine-cosine algorithm, without jointly handling IRS reconfiguration, technology selection and power allocation. The problem is given in full as Eq. (12) with constraints (12a)-(12g), and the sources of non-convexity are itemized: binary selection variables, their bilinear coupling with power variables, and the nonlinear dependence of the VLC rate on mirror orientation.

Reformulates the problem as an MDP and designs a two-agent DRL framework under centralized training with decentralized execution: a selection-and-power agent handles scheduling variables and an IRS-orientation agent handles mirror angles, ordered so the mirror is reconfigured first and selection follows, with both agents sharing one cooperative reward. DQN-style methods handle only discrete actions and would require quantizing continuous power fractions and angles, causing combinatorial explosion; DDPG-style single-agent methods must learn binary selection, power and angles within one heterogeneous action space, coupling variables of different scales. The state space, the two action spaces (Eqs. 13 and 14), the cooperative reward with QoS and power-budget penalty terms (Eq. 15), and the training procedure (Algorithm 1) are specified; training complexity is analyzed and per-slot execution reduces to one forward pass per actor, independent of the channel realization.

Simulations show the framework improves both sum rate and fairness over baselines: it converges after about 2000 episodes to 350 Mbps, within 10% of the 392 Mbps exhaustive-search bound, exceeds DDPG and DQN by 13% and 25%, and oscillates by less than 5 Mbps. SCA reaches 330 Mbps but must be re-solved at every scheduling slot; DQN saturates early at 280 Mbps because of its discrete action space; DDPG reaches a higher steady-state rate but oscillates by about 25 Mbps around the SCA benchmark. Based on training curves comparing algorithms under Python 3.12 and PyTorch 2.0 over 10000 episodes of 100 time steps, reporting convergence values and oscillation amplitudes.

Ablation and fairness analysis quantify the IRS agent's contribution: sum rate first rises then falls with user count, peaking at 415 Mbps for the proposed method; disabling the IRS agent lowers the peak to 380 Mbps and fairness to 0.81, while the full framework stays above 0.93 at high user density. Among single-subnetwork baselines, VLC only beats RF only because of its larger bandwidth, yet both fall below hybrid schemes, indicating technology diversity is not replaceable for LoS-blocked users. Uses Jain's fairness index to compare the proposed method, SCA, the IRS-agent-disabled variant, VLC only and RF only across user counts, and explains why blocked users cannot be recovered by a single subnetwork.

Perspective

The results target indoor downlink IRS-assisted hybrid RF/VLC scenarios, for mobile users equipped with multi-branch angle diversity receivers and RF antennas and coordinated by a central controller, moving under a random waypoint model with LoS blockage modeled as an independent Bernoulli process. The method's value lies in using an offline-trained policy for online execution: training is offline, and per-slot execution needs only one forward pass per actor, making it suited to real-time scheduling where the channel keeps changing with user mobility and blockage. For readers reusing the idea, the directly transferable parts are the action-space decomposition, assigning scheduling-type variables and physical-orientation-type variables to different agents while keeping coordination through a shared reward, and encoding constraint violations into the reward.

All results are simulation-based, with no hardware prototype or measured channel, so the effects of mirror orientation control precision, MEMS rotation latency and real blockage statistics remain open questions. The fairness and rate conclusions rest on a specific indoor size, AP and IRS array configuration, user speed and blockage probability distribution; behavior under changed parameters needs separate validation. The proposed method still trails the exhaustive-search bound by about 10%, and sum rate rises then falls with user count, with the peak position tied to interference and the difficulty of meeting the QoS constraint, so behavior at larger scale or under different QoS requirements merits further observation. In addition, centralized training requires global channel state information while decentralized execution relies only on local observations, leaving the quality of coordination under limited observability a direction worth continued study.

Sources