Skip to main content
Back to timeline
NVIDIA Technical BlogSource publication:

NVIDIA releases AIPerf: a multiprocess load client succeeding GenAI-Perf, demonstrated with Qwen3-0.6B on static and Poisson LLM inference benchmarks

Synopsis

NVIDIA introduces AIPerf, the designated successor to GenAI-Perf and a ground-up rewrite that uses a multiprocess architecture coordinated over ZMQ to keep the client from becoming the bottleneck, supports 15+ endpoint types and constant, Poisson, and gamma arrival patterns, and is demonstrated with Qwen3-0.6B served through vLLM in two measurement loops: a static benchmark pinned to 128 input and 128 output tokens, and a Poisson benchmark at an average of 10 requests per second with input lengths of 512 plus or minus 128.

AI-generated editorial illustration: Benchmarking LLM Inference at Scale with AIPerf

Interpretation

AIPerf is the designated successor to GenAI-Perf and a ground-up rewrite: it does not run on top of Perf Analyzer the way GenAI-Perf did, and instead uses a multiprocessed system in which worker processes generate load, separate record-processor services handle results, and everything is coordinated over ZMQ. The text presents this architectural break as the reason AIPerf can scale, and notes that most benchmarkers, GenAI-Perf included, use a single-process architecture that becomes GIL-bound under real concurrency or request rate. The evidence is the text's direct statement of the design, which is a tool-design description rather than a controlled experiment; no quantitative AIPerf-versus-GenAI-Perf comparison is given.

AIPerf supports 15+ endpoint types (chat, responses, NIM rankings, image generation, and more) along with public datasets such as ShareGPT and trace replay formats from Mooncake, Baseten, WEKA (AgentX), and others. The text emphasizes that a quick synthetic smoke test and replaying captured production traffic do not require a different tool, and that load shape is controllable through constant, Poisson, and gamma arrival patterns with tunable burstiness, gradual ramping for concurrency and request rate, and synthetic distributions including vLLM/SGLang range-ratio. The evidence is the text's enumeration of supported features; no validation results for individual endpoint types or datasets are provided.

The text demonstrates two commands against Qwen3-0.6B served through vLLM: a static benchmark pinning both input and output to exactly 128 tokens (standard deviations of 0, with min_tokens:128 and ignore_eos:true forcing the output length), and a Poisson benchmark at an average of 10 requests per second with exponentially distributed inter-arrival times, input of 512 plus or minus 128, output of 128 plus or minus 32, random seed 42, and 200 requests. The text states that the static command reproduces a commonly used fixed-length benchmark, while the Poisson command introduces bursts and gaps and variable prefill lengths closer to real queuing, and notes that dropping min_tokens and ignore_eos deliberately releases the output-length constraint. The evidence is the full command lines and parameter explanations given in the text; the text reports that input sequence lengths in the Poisson run ranged from 154 to 818 tokens and that the run's metric distributions were noticeably wider than the static baseline.

AIPerf reports four core metrics, TTFT, ITL, request latency, and output token throughput, each with percentile breakdowns (p25, p50, p75, p90, p95, p99) alongside minimums, maximums, averages, and standard deviations; when DCGM or pynvml is available it also pulls GPU power draw, utilization, and memory consumption into the same run output. The text stresses that percentile breakdowns can highlight long-tail distributions, since a server with a healthy mean TTFT and an outlier p99 looks fine in aggregate and fails in production, and that correlating a latency spike with a memory pressure event does not require a separate profiling session. The evidence is the text's description of metric definitions and outputs; the text notes that --streaming is required to measure TTFT and ITL, because without it the server batches the full response before sending.

Perspective

The walkthrough targets readers serving Qwen3-0.6B through vLLM on a single GPU, with the stated goal of establishing the measurement loop rather than benchmarking that model; the text explicitly says swapping in a different model or endpoint is a one-flag change. The approach suits inference teams that need to separate prefill from decode behavior and watch long-tail latency and GPU telemetry, and the text claims the same tool extends to multi-node Kubernetes deployments, KV cache reuse warm-up, production trace replay, prefix synthesis, custom datasets, and sweep configurations. The text also notes an aarch64 platform caveat: the crick dependency ships as source-only and requires a C toolchain (build-essential on Debian/Ubuntu, Development Tools on RHEL).

The text presents the static and Poisson runs through Figures 2, 4, 5, and 6, but the prose gives only qualitative descriptions (the Poisson run's distributions are noticeably wider, TTFT spread is larger, input lengths range from 154 to 818 tokens) without listing specific values, so the magnitude of the differences cannot be checked from the text alone. The text states that the multiprocess architecture prevents AIPerf from becoming a client-side bottleneck but gives no measurement of client-side resource use or saturation point. The 15+ endpoint types, trace replay formats, and complex scenarios (multi-node Kubernetes, KV cache warm-up, sweep configurations) are listed but not demonstrated here. The text also does not state the GPU model, vLLM version, or AIPerf version used for the Qwen3-0.6B demonstration, so reproduction requires confirming the environment independently.

Sources