Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Beaver shares an H200 GPU between vRAN and Llama-3.3-70B, holding the 1.5ms uplink deadline while keeping 74% of serving throughput

Beaver is a GPU sharing system for the AI-RAN setting that sizes the vRAN's streaming multiprocessor (SM) allocation from each slot's scheduled workload, repartitions SM allocations at slot granularity, and rewrites compiled ML kernels to yield HBM bandwidth during the vRAN's memory-critical phases; evaluated on an H200 with NVIDIA Aerial, heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving, it keeps the vRAN's p99.9 latency within its 1.5ms uplink deadline while retaining 74% of Llama-3.3-70B serving throughput, incurs no observed deadline misses under replayed cellular traces, protects a 375us downlink deadline, and generalizes to A100, GB10 and GH200.