NVIDIA ran all 1,007 turns of the MLPerf Edge Agentic benchmark on a single Jetson AGX Thor with TensorRT Edge-LLM in 24 minutes 36 seconds, 6.4x faster than the llama.cpp reference
Synopsis
In the MLPerf Inference v6.1 Edge Agentic benchmark, NVIDIA's TensorRT Edge-LLM ran Qwen3.6-27B in SingleStream mode on a single Jetson AGX Thor Developer Kit with 128 GB of unified memory at MAXN power mode, reaching 52.33 tokens per second output throughput, 247.12 ms median time to first token, 14.68 ms median time per output token, and 87.94% BFCL overall accuracy, and completing all 1,007 turns of the performance workload in 24 minutes 36 seconds, 6.4x faster than the llama.cpp reference submission's 2 hours 37 minutes, using NVFP4 quantization, FP8 KV cache, KV cache and recurrent-state reuse, and tree-based multi-token prediction.
Interpretation
The submission completed the MLPerf Edge Agentic performance phase of 20 conversations and 1,007 generated turns on a single Jetson AGX Thor in 24 minutes 36 seconds, with 52.33 tokens per second output throughput, 247.12 ms median time to first token, and 14.68 ms median time per output token. The MLCommons Edge Agentic example also publishes a llama.cpp reference run on the same hardware using Qwen3.6-27B with Q4_K_M quantization, which completes in 2 hours 37 minutes; this submission cuts the same workload to one 6.4th of that time. Results come from the official MLPerf Inference v6.1 Edge Agentic measurement flow; the performance phase replays recorded software-engineering agent trajectories and also measures IoU-based inline accuracy to confirm the agent is running correctly, with temperature 0, seed 42, reasoning disabled, and concurrency 1.
The accuracy phase uses BFCL v4 prompts with single-turn only and reasoning off, where the submission reached 87.94% overall accuracy, evaluating whether the model selects the correct function, generates valid arguments, and avoids calling a tool when none is required. Quantization and speculative decoding usually trade accuracy for speed; this submission reports the accuracy level MLPerf requires while using NVFP4 weights and activations, FP8 KV cache, and tree-based MTP. Accuracy is measured with BFCL v4 prompts on edge devices under a single-turn, reasoning-off setting, which the text describes as balancing accuracy and evaluation time on edge devices.
The runtime identifies reusable prompt prefixes and restores their cached attention KV pages; because Qwen3.6 uses a hybrid architecture, the runtime also restores the recurrent state and partial KV-page state required to continue execution correctly, then prefills only the new suffix of the conversation. The work extends KV cache and recurrent-state reuse to agent trajectories on a hybrid architecture, making cache reuse complementary to tree-based MTP: the former reduces cost before generation begins, the latter reduces target-model steps during generation. On this workload about 96% of prompt tokens are served with hot cache, and the runtime prefills only about 0.5M of the total 13.6M prompt tokens across the turns.
Beyond traditional linear MTP, TensorRT Edge-LLM implements tree-based MTP: high-probability candidates are organized into a tree, the target model verifies candidates in one forward pass and the runtime accepts the matching path, advancing generation by several tokens when multiple candidates are accepted. The MLPerf server configuration uses 8 draft steps, the top-2 candidates at each drafting depth, and a 16-node verification tree; compared with a linear MTP with 3 draft steps, tree-based MTP could achieve an additional ~40% decoding performance gain for this workload. The text notes tree-based verification is useful for function calling because tool names, JSON syntax, and common argument structures are often predictable, while multiple branches can preserve likely alternatives for individual argument values; drafting parameters can be adjusted when launching the server.
Perspective
The result targets teams deploying an OpenAI-compatible model endpoint on a single Jetson AGX Thor Developer Kit (128 GB unified memory, MAXN power mode, SingleStream, concurrency 1), for edge agent scenarios that need long context and tool calling. It lets developers start from a published calibrated Qwen3.6-27B NVFP4 checkpoint, or perform post-training quantization once on any development system before deploying, and reproduce the submission using the export settings, TensorRT engine build commands, server configuration, and MLPerf client configuration in the release branch.
The accuracy phase uses BFCL v4 with single-turn only and reasoning off, a trade-off made to balance accuracy and evaluation time on edge devices, so accuracy under multi-turn tool calling remains an open question. The ~40% additional decoding gain from tree-based MTP is a comparison against a linear MTP with 3 draft steps on this workload, and behavior on other workloads and drafting parameters remains to be seen. Also, while this is the full text, it does not report power, energy efficiency, or comparisons with cloud deployment, which remain to be filled in.
