Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN
Synopsis
This work demonstrates on a live O-RAN testbed that two autonomous agents with individually correct objectives—one protecting a latency SLA and one maximizing utilization for energy efficiency—jointly drive recurring opposing excursions of the shared resource partition, and presents and proves a lightweight arbitration layer, AURA, whose three admission checks (feasibility invariants, per-variable dwell time, deadband) reduce shared-state excursion amplitude from 8.4 to 0.4 PRB and cross-slice throughput starvation from 40–55% to 0.3%, while leaving the protected slice's own latency compliance unchanged.
Fig. 1: AURA architecture. Autonomous AI agents ( A sl A_{\mathrm{sl}} , A en A_{\mathrm{en}} ) in the Non-RT RIC submit proposals to a central arbiter, which enforces three conditions (feasibility invariant, dwell time, deadband) with SLA-restoring actions taking priority. Admitted proposals update a shared quota table reloaded by the OAI gNB; the NR MAC downlink pre-processor enforces per-slice PRB quotas. One-way latency ( ℓ 1 \ell_{1} ) is measured via a dedicated UDP probe to UE 1 ; slice-2 throughput ( r 2 r_{2} ) is read at UE 2 ; MAC statistics return via E2 to FlexRIC and feed a shared telemetry store read by both agents.
arXivInterpretation
First end-to-end empirical demonstration of multi-agent RAN instability on a running 5GSA O-RAN stack: two agents that are each correct alone jointly drive recurring excursions of the shared resource partition. Prior conflict-mitigation work assumes a statically known application population analyzable offline; this work shows runtime-emergent agent conflicts validated with real measured one-way delay and throughput rather than model-derived proxies. Experiments on an OAI/FlexRIC testbed with real measured one-way delay and throughput, 5 repetitions per regime (20 runs total); under Direct the settled C amplitude is 8.4 PRBs, with 2.0 PRBs average movement in the second independent load cycle and at least 2 PRBs in 3 of 5 repetitions (per-repetition values 0, 2, 0, 2, 6); amplitude is zero when either agent runs alone.
Formalizes the pathology as a delayed opposing best-response process, with a limit-cycle proof (Proposition 1) and a convergence guarantee for the arbitrated system (Proposition 2). Maps classical tools such as dwell-time conditions from switched-system stabilization onto the O-RAN control plane, identifying the shared-variable structure, delay sources, and an arbitration rule whose admitted-action sequence provably terminates. Proposition 1 gives a limit-cycle amplitude lower bound of δ·⌈τ_loop/T_a⌉; Proposition 2 proves the finite improvement property and termination under τ_d>τ_loop and 0<ε≤min(δ,δ_c); the experiment sets τ_d=8 s above the measured τ_loop of about 7 s and enforces this as a precondition of the experiment.
Implements per-slice PRB quota enforcement in the OAI NR MAC downlink pre-processor, absent in upstream OAI, and measures the full control-loop latency budget from agent decision to MAC enforcement. Documents that the current upstream OAI E2 slice service model is an emulator (indication messages filled with random data, control messages acknowledged without touching the MAC), so multi-agent studies on that path without verifying enforcement measure nothing; this work implements enforcement directly and rewires the E2 callbacks to the same quota table. In a single-UE sweep, raising quota from 13 to 38 PRBs raises measured receiver goodput from 17.0 to 43.1 Mbps (a 2.54× increase consistent with the quota ratio), while the MAC per-slot allocation counter tracks the configured quota exactly; arbiter admission logic overhead is negligible relative to the roughly 265 ms enforcement step.
AURA trades an explicitly acknowledged cost for system-wide feasibility: it strongly damps shared-state excursions and virtually eliminates cross-slice throughput starvation, but does not improve the protected slice's own latency compliance. Reveals that no single-slice SLA metric captures the cross-slice cost: Direct's nominally lower slice-1 violation rate (84.5% versus AURA's 92.9%) is achieved by starving slice 2 by 40–55% in throughput, not a genuine gain. AURA reduces slice-2 throughput violations from 40–55% under Direct and Single to 0.3%; a static control run at AURA's settled operating point (q1,q2,C)=(26,18,44) shows the gap to Static on slice-1 latency shrinks to 4.2 percentage points but is not eliminated.
Perspective
The result targets autonomous controllers with contradictory objectives sharing delayed-observation loops over a common resource, applicable to multi-second-horizon slice-quota and energy-cap decisions while leaving sub-second scheduling to the near-RT RIC. It enables operators and RIC developers to place a provably correct admission boundary before actions reach the network, and because that boundary is agnostic to agent internals it applies equally to rule-based and learning-based agents. The authors note the framework carries over to scenarios such as power control versus interference management and traffic steering versus energy saving, and plan to extend evaluation to multi-UE, multi-cell, many-agent over-the-air testbeds.
The experiment fixes T_a=5 s and does not sweep T_a or step sizes (δ,δ_c) to test the amplitude scaling predicted by Proposition 1, nor does it quantify how conflict rates scale with agent population; the authors explicitly do not claim platform-scale concurrent-agent counts. The agents are rule-based, so the arbiter's guarantees for learned policies remain to be validated with a DRL or LLM-backed planner on an over-the-air testbed. Enforcement is file-driven, and the roughly 265 ms enforcement step is the τ_loop bottleneck; a production E2-based path would need this component re-measured. The residual 4.2-percentage-point gap between slice-1 latency and Static is attributed to live-dynamics effects not present in the static control, with precise attribution left to future work.
