Public articles linked to the same research event.
arXiv Beaver is a GPU sharing system for the AI-RAN setting that sizes the vRAN's streaming multiprocessor (SM) allocation from each slot's scheduled workload, repartitions SM allocations at slot granularity, and rewrites compiled ML kernels to yield HBM bandwidth during the vRAN's memory-critical phases; evaluated on an H200 with NVIDIA Aerial, heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving, it keeps the vRAN's p99.9 latency within its 1.5ms uplink deadline while retaining 74% of Llama-3.3-70B serving throughput, incurs no observed deadline misses under replayed cellular traces, protects a 375us downlink deadline, and generalizes to A100, GB10 and GH200.
Beaver is a GPU sharing system for the AI-RAN setting that sizes the vRAN's streaming multiprocessor (SM) allocation from each slot's scheduled workload, repartitions SM allocations at slot granularity, and rewrites compiled ML kernels to yield HBM bandwidth during the vRAN's memory-critical phases; evaluated on an H200 with NVIDIA Aerial, heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving, it keeps the vRAN's p99.9 latency within its 1.5ms uplink deadline while retaining 74% of Llama-3.3-70B serving throughput, incurs no observed deadline misses under replayed cellular traces, protects a 375us downlink deadline, and generalizes to A100, GB10 and GH200.
Beaver is a GPU sharing system for the AI-RAN setting that sizes the vRAN's streaming multiprocessor (SM) allocation from each slot's scheduled workload, repartitions SM allocations at slot granularity, and rewrites compiled ML kernels to yield HBM bandwidth during the vRAN's memory-critical phases; evaluated on an H200 with NVIDIA Aerial, heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving, it keeps the vRAN's p99.9 latency within its 1.5ms uplink deadline while retaining 74% of Llama-3.3-70B serving throughput, incurs no observed deadline misses under replayed cellular traces, protects a 375us downlink deadline, and generalizes to A100, GB10 and GH200.
Beaver is a GPU sharing system for the AI-RAN setting that sizes the vRAN's streaming multiprocessor (SM) allocation from each slot's scheduled workload, repartitions SM allocations at slot granularity, and rewrites compiled ML kernels to yield HBM bandwidth during the vRAN's memory-critical phases; evaluated on an H200 with NVIDIA Aerial, heterogeneous multi-cell workloads, real-world cellular traces, production inference kernels, and full-stack LLM serving, it keeps the vRAN's p99.9 latency within its 1.5ms uplink deadline while retaining 74% of Llama-3.3-70B serving throughput, incurs no observed deadline misses under replayed cellular traces, protects a 375us downlink deadline, and generalizes to A100, GB10 and GH200.