Skip to main content
Back to timeline
arXivSource publication:

Beaver shares an H200 GPU between vRAN and Llama-3.3-70B, holding the 1.5ms uplink deadline while keeping 74% of serving throughput

Related research and updates

Synopsis

Beaver is a GPU sharing system for the AI-RAN setting that sizes the vRAN's streaming multiprocessor (SM) allocation from each slot's scheduled workload, repartitions SM allocations at slot granularity, and rewrites compiled ML kernels to yield HBM bandwidth during the vRAN's memory-critical phases; evaluated on an H200 with NVIDIA Aerial, heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving, it keeps the vRAN's p99.9 latency within its 1.5ms uplink deadline while retaining 74% of Llama-3.3-70B serving throughput, incurs no observed deadline misses under replayed cellular traces, protects a 375us downlink deadline, and generalizes to A100, GB10 and GH200.

Source-provided article image: Beaver: Elastic GPU Sharing between ML and Latency-Critical vRAN Workloads
Figure 1 ·

Figure 1: AI-and-RAN co-tenancy setting. Radio units deliver

arXiv · Page 1

Interpretation

Beaver presents a GPU sharing system that jointly manages compute and memory resources to co-locate latency-critical vRAN workloads with best-effort ML workloads on shared GPUs while protecting the vRAN's strict processing deadline. The work targets the inherent asymmetry of co-location in AI-and-RAN: vRAN is latency-critical while ML is a throughput-oriented, best-effort co-tenant, and Beaver brings both compute and memory into one sharing management rather than addressing only one side. The abstract states the system was implemented and evaluated using NVIDIA Aerial with heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving.

Beaver sizes the vRAN's SM allocation from each slot's scheduled workload and repartitions SM allocations at slot granularity. The allocation granularity is aligned with the vRAN's slot scheduling cadence, so SM resources follow each slot's scheduled workload rather than a static or coarse partition. The abstract lists this mechanism as a core component of the system and evaluates on an H200 using whether p99.9 latency stays within the 1.5ms uplink deadline.

Beaver rewrites compiled ML kernels so they yield HBM bandwidth during the vRAN's memory-critical phases. This treats memory bandwidth as a shareable resource that can be yielded, sitting alongside the SM allocation mechanism as part of joint compute-and-memory management. The abstract lists it among the system mechanisms and reports no observed deadline misses under replayed cellular traces and protection of a 375us downlink deadline.

On an H200, Beaver keeps the vRAN's p99.9 latency within its 1.5ms uplink deadline while retaining 74% of Llama-3.3-70B serving throughput, and generalizes to A100, GB10 and GH200. The result reports quantified behavior on both sides at once, latency protection and ML throughput retention, and reports generalization across several GPUs. The abstract reports the specific values of 74% throughput retention, a 1.5ms uplink deadline, a 375us downlink deadline, and generalization to A100, GB10 and GH200; these are abstract-level reported values.

Perspective

The work targets the AI-and-RAN co-location setting within AI-RAN: running virtualized radio access network (vRAN) workloads and AI services on the same GPU, where vRAN is the latency-critical party and ML is a throughput-oriented, best-effort co-tenant. Beaver's premise is that the vRAN has a slot-organized scheduled workload, so SM allocation can be sized from each slot's scheduled workload and repartitioned at slot granularity; in parallel, ML-side kernels can be rewritten to yield HBM bandwidth during the vRAN's memory-critical phases. The abstract reports protection targets of a 1.5ms uplink deadline and a 375us downlink deadline, with evaluation on an H200 using NVIDIA Aerial, heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving, and states generalization to A100, GB10 and GH200. For engineering teams wanting to layer latency-critical wireless processing and AI services on one accelerator, this offers a reproducible mechanism direction: adjust SMs on the slot cadence and yield bandwidth in memory-critical windows.

The reading scope here is the abstract only, so the specific SM allocation algorithm, the implementation of the kernel rewriting, the overhead of slot-granularity repartitioning, and the full experimental setup behind the 74% throughput retention and the deadline values (workload scale, trace length, comparison conditions) are not developed in the text and would need the original. The abstract reports p99.9 latency and no observed deadline misses under replayed traces; how such tail metrics behave over longer traces, more cells, or different ML workload mixes remains an open question. The abstract also states generalization to A100, GB10 and GH200 but does not give corresponding values per platform, so cross-platform consistency awaits the original data.

Sources