Skip to main content
Back to timeline
NVIDIA Technical BlogSource publication:

A multi-agent skill ported all 24 TileGym operators from cuTile Python/Triton-TileIR to cuTile Rust, averaging 99.5% of reference performance

Synopsis

The work built a multi-agent AI skill that translates GPU tile kernels written in cuTile Python and Triton-TileIR into cuTile Rust, and used it to port all 24 public TileGym operators (roughly 40 kernels, spanning element-wise operations through flash-attention decode, MLA, and MoE), measuring a geometric mean of 0.995 by CUPTI device time on NVIDIA DGX B200, i.e. 99.5% of cuTile Python performance on average, with every stage gated by machine-checkable verdicts and Tile IR diffs.

AI-generated editorial illustration: Translating CUDA Tile Operations from Python to Rust Using Agentic AI

Interpretation

It proposes and ships a multi-agent translation pipeline that converts cuTile Python and Triton-TileIR kernels into cuTile Rust, covering analysis, the device kernel, host and FFI code, and benchmarking. Cross-front-end porting previously meant manual rewriting; here the flow is split into specialized subagents while the top-level agent only routes, and each stage loads only the reference documents it needs. The text shows the skill layout (SKILL.md, agents/*.md, references/, scripts/, examples/) and states that every subagent return must end with a literal block plus one VERDICT line, on which the orchestrator routes.

It uses the shared CUDA Tile IR (cuda_tile dialect) as the verification backbone, applying Tile IR diff structurally before any test runs. Functional tests can pass a wrong-but-plausible translation (for example a TMA load with a wrong cost hint or a dropped divisibility attribute); IR diff directly compares memory-op families, tile shapes, and reductions, surfacing such issues early. The softmax example shows Rust compiling to the same op inventory as Python: one view load, reduce_max/reduce_sum on the right axis, one view store, and TMA on both ends; IR diff serves both as the kernel writer's self-check and as the IR-diff analyst's deep comparison.

Using the skill, all 24 public TileGym operators were ported, roughly 40 GPU kernels, averaging 99.5% of cuTile Python performance. The numbers come from the CI benchmark pipeline itself: CUPTI device time on NVIDIA DGX B200 with one exclusive GPU per backend, 347 paired configurations across the 24 operators, taking the best measurement across four CI runs per configuration. Overall geomean is 0.995, all 24 operators clear the 0.95 check, about a third come out ahead of the reference, with the largest wins on element-wise and normalization kernels; each conversion lands as a standard six-file changeset.

It locates the translation difficulty in implicit versus explicit specialization and gives concrete correspondences. cuTile Python's JIT specializes implicitly at call time and drops untaken ct.Constant branches; cuTile Rust requires every specialization in the signature, so one Python kernel can become multiple Rust entries (layer_norm splits into 2-Dnchw and 1-Dw1 because the branch changes tile rank). The softmax line-by-line comparison shows the mapping: Constant[int] becomes a const generic, ct.load with NEG_INF padding becomes make_partition_view plus Partition::load, ct.bid(0) becomes get_tile_block_id(), and a keepdims reduction becomes reduce_* plus explicit reshape and broadcast.

Perspective

The result targets teams using the CUDA Tile family (cuTile Python, Triton-TileIR, cuTile Rust) who want to migrate kernel libraries to Rust, provided the environment meets CUDA 13.1+, a Blackwell GPU for the perf check, Rust 1.89+, and the tileiras compiler. The skill ships with the TileGym repo at skills/tilegym-converting-python-to-rust/, and the kernels live under src/tilegym/ops/cutile_rs/; a user can point an agent at the repo and ask it to add a cutile-rs backend, with the pipeline handling analysis, kernel, FFI, correctness, and benchmarking. The performance conclusion applies to operator-to-operator comparison under CUPTI device time; the authors state they measure wall clock but report device time to keep the comparison about the kernels themselves.

The authors note that not all shipped TileGym kernels yet use the fully safe style shown: because each port must reproduce the reference kernel Tile IR exactly, where only an unchecked API reproduces it the port uses that API, and the team is still migrating those kernels onto the safe surface. Also, cuTile Rust can emit Tile IR directly, so in principle one could write a kernel matching other front ends exactly, but such kernels become uninterpretable, so the skill is biased toward idiomatic code; the authors expect that matching the emitted Tile IR exactly would match performance exactly, though the experiments capture only device time. Readers should also note that the performance numbers come from the CI benchmark pipeline and depend on specific hardware and toolchain versions; the text mentions a full conversion runs on the order of millions of tokens and that the orchestrator has hard spawn caps, so a run either converges within budget or stops with a diagnosis on disk.

Sources