Engineering

Helion PyTorch Kernel DSL: Meta's PyTorch-Native GPU Compiler (2026)

Back to BlogWritten by Published Sep 30, 2026Updated
Helion PyTorch KernelHelion vs TritonPyTorch Native Kernel CompilerHelion AutotuningGPU Kernel DevelopmentTorchInductorGPU CloudH100AMD MI300X
Helion PyTorch Kernel DSL: Meta's PyTorch-Native GPU Compiler (2026)

A Helion PyTorch kernel starts as ordinary Python and compiles straight into autotuned Triton code, and Meta's own numbers show why kernel authors are paying attention: Meta's attention kernel example runs 30 lines in Helion against 120 lines of hand-written Triton and "thousands of lines" of raw CUDA, a 4x cut in source you have to maintain by hand. That's not a marketing rounding error, it's the literal ratio in Meta's own benchmark writeup. This post walks through what Helion actually is, how it relates to torch.compile and TorchInductor (the tools most people researching "pytorch native kernel compiler" already know), and what the autotuning search space looks like in practice. The performance figures below are Meta's own published numbers, read closely and checked against the autotuner's own config reference rather than an independent lab rerun, and the caveats that come with a vendor's own benchmark are called out where they matter, including whether the productivity story holds up once you factor in a real multi-vendor GPU fleet.

TL;DR: What Is Helion, Meta's PyTorch-Native Kernel DSL?

  • What it is: Helion is a Python-embedded DSL that compiles into autotuned Triton kernels, released in Beta by Meta on Oct. 22, 2025.
  • The line-count claim: Meta's own attention kernel example is 30 lines in Helion versus 120 lines in hand-written Triton and thousands in CUDA, a roughly 4x cut.
  • Autotuning cost: A first, uncached run typically takes around 10 minutes and searches thousands of Triton configs before locking in a config for later runs.
  • Backend reach: Helion compiles to Triton, so NVIDIA CUDA and AMD ROCm are both live paths; Intel GPUs run through Triton's separate community-maintained Intel backend rather than anything Helion documents directly.
  • On Spheron: you can validate a Helion kernel's autotuned config on $2.65/hr NVIDIA H100 and $3.59/hr AMD MI300X instances (as of 17 Sep 2026), spinning up either in under 2 minutes. For another kernel-level benchmark on the same hardware family, see Spheron's FlashAttention-4 on Blackwell guide.

What Is Helion? A PyTorch-Native Kernel Compiler Explained

A PyTorch-native kernel compiler takes code written in PyTorch's own idiom, rather than a separate systems language, and lowers it directly to a fast GPU kernel. Helion is Meta's entry in that category: it "resolves this conflict by compiling a high-level Python-embedded domain-specific language (DSL) into automatically tuned Triton code," as Meta describes it, where "this conflict" is the tradeoff between writing fast, hardware-specific kernels and writing portable, maintainable ones. You write ordinary PyTorch tensor operations inside a tiled loop, and Helion's compiler turns tiling decisions, memory layout, and loop order into a search space that gets autotuned per GPU instead of hand-coded once.

The programming model is described internally as "PyTorch with Tiles." Code outside the outermost hl.tile loop is standard host-side PyTorch, used for shape setup and output allocation. Code inside that loop is device code: it gets compiled into a single Triton kernel that runs on the GPU. The one core construct that separates a Helion kernel from a normal PyTorch function is hl.tile, which subdivides an iteration space into tiles without specifying their size, order, or memory layout. Those three decisions become the autotuner's job.

This matters for the "helion gpu dsl" search intent specifically: Helion is not a new low-level language to learn. If you already know PyTorch tensor operators (torch.addmm, elementwise ops, reductions), Meta's own framing is direct: "Familiarity with PyTorch means you already know most of Helion." The DSL surface is small; the leverage comes from what the compiler does with it, and the same tile-and-fuse logic underpins why kernel fusion cuts inference cost in the first place: fewer round-trips to HBM per operation.

Because Helion's output is Triton, not a new instruction set, it helps to have Triton fundamentals first. If tile-level GPU programming in Python is new to you, our OpenAI Triton kernel development guide covers the concepts Helion builds on top of; this post assumes you're comfortable with the idea of writing GPU kernels in Python at all, and focuses on what Helion automates that Triton leaves manual.

Helion vs torch.compile and TorchInductor: Where It Actually Sits

Most people researching "pytorch native kernel compiler" today land on torch.compile and TorchInductor content, and for good reason: TorchInductor is PyTorch's default compiler backend, and it's also a component Helion depends on directly. The relationship isn't competitive, it's layered.

torch.compile optimizes an existing PyTorch model automatically. You add a decorator, and TorchInductor traces the model's graph, fuses operators where it can, and generates Triton (or C++) code without you writing any kernel logic yourself. It's aimed at getting reasonable speedups on code you didn't design for a compiler.

Helion is for the opposite situation: kernels you deliberately author because the default fusion and codegen TorchInductor produces for a full model graph isn't tight enough for a specific hot operation. Inside a Helion kernel body, you still use standard PyTorch operators, and Meta's writeup is explicit that "Helion leverages TorchInductor, a core component of PyTorch 2, to automatically map these PyTorch calls to their corresponding low-level Triton implementations." TorchInductor is the codegen backend Helion's compiler pipeline calls into during the final lowering stage; Helion adds the tile abstraction, the explicit kernel boundary, and, most importantly, the autotuning engine that TorchInductor's whole-graph fusion doesn't search as exhaustively for a single custom kernel.

Put simply: torch.compile/TorchInductor optimizes what you already wrote. Helion is what you reach for when you're writing a new kernel on purpose and want compiler-driven tuning instead of hand-placed configs.

Why Helion Lives Under the PyTorch Foundation, Not Just Meta

Helion was built at Meta and released open source, but it didn't stay a single-vendor side project. It now sits in the pytorch GitHub organization, and the PyTorch Foundation states plainly that it is "the vendor-neutral home for the open source intelligence layer developers use for training, optimizing, serving, orchestrating, and running models on any chip in any cloud for any agent," and that, "hosted by the Linux Foundation, the PyTorch Foundation supports the core PyTorch framework alongside a growing portfolio of innovative projects including vLLM, DeepSpeed, Ray, Helion, and Safetensors."

That governance detail is not just paperwork. Meta's writeup lists Helion's own acknowledgements as contributors from "Meta, NVIDIA, AMD, and Intel," which lines up with the vendor-neutral framing: a kernel DSL that's supposed to target every major GPU vendor is more credible with commit access spread across those vendors' engineers rather than gated behind one company's roadmap. It's also why the multi-backend story later in this post (NVIDIA, AMD, and an emerging Intel path) isn't just an aspiration in a README, it's baked into who's actually shipping the compiler.

Helion vs Triton vs CUTLASS: The Control-vs-Productivity Tradeoff

Kernel authoring tools sit on a spectrum between raw control and developer productivity, and Helion, Triton, and CUTLASS's CuTe DSL each occupy a different point on it.

ToolAbstraction levelWhat you specify by handWhat gets searched automatically
CUDA / CUTLASS CuTe DSLLowest (tile/thread level)Memory layout, tiling, pipelining stages, warp specializationLittle to nothing without extra tooling
TritonMid (block level)Block sizes, loop structure, memory access pattern, autotune config listOnly the configs you explicitly enumerate in triton.autotune
HelionHigh (tile level, PyTorch syntax)The tiling boundary (hl.tile) and the operators inside itBlock sizes, loop order, flattening, indexing mode, reduction strategy, PID mapping

Meta frames this directly: "While Triton represents a major step forward, it still requires significant manual effort. Developers are responsible for explicitly managing tensor indexing, defining search spaces for autotuning, managing kernel arguments, and changing the optimization strategy can often require significant code rewrites." That last point is the practical difference. In Triton, switching from pointer-arithmetic indexing to block pointers or tensor descriptors is a rewrite. In Helion, it's a value in a config the autotuner already tries.

CUTLASS's CuTe DSL sits below Triton on this spectrum, and our CuTe DSL guide covers that end of the tradeoff in more depth: more manual control over Blackwell-specific features like Tensor Memory Accelerators, correspondingly more code to write and tune by hand. Helion doesn't try to compete at that level of control. Its bet is that for most kernels, giving up some manual placement in exchange for a compiler-driven search across thousands of configurations wins on wall-clock developer time without giving up much run-time performance, a bet Meta's own benchmarks below try to substantiate.

Writing and Autotuning a Real Helion PyTorch Kernel (Code Walkthrough)

Here's Meta's own matmul kernel, reproduced exactly as published, to show what the abstraction level actually looks like in practice:

python
import torch, helion, helion.language as hl

@helion.kernel()
def matmul(x: torch.Tensor, y: torch.Tensor) -> torch.Tensor:
    # --- Host Code (runs on CPU) ---
    m, k = x.size()
    k, n = y.size()
    out = torch.empty([m, n], dtype=x.dtype, device=x.device)

    # --- Device Code (compiles to a Triton kernel) ---
    for tile_m, tile_n in hl.tile([m, n]):
        acc = hl.zeros([tile_m, tile_n], dtype=torch.float32)
        for tile_k in hl.tile(k):
            acc = torch.addmm(acc, x[tile_m, tile_k], y[tile_k, tile_n])
        out[tile_m, tile_n] = acc

    return out

Everything above the hl.tile loop is ordinary Python and PyTorch: shape reads, an output allocation. Everything inside the nested hl.tile loops is what gets compiled into a Triton kernel. Notice what's absent compared to a hand-written Triton matmul: no explicit BLOCK_SIZE_M/BLOCK_SIZE_N/BLOCK_SIZE_K constants, no manual pointer-stride arithmetic, no triton.autotune config list to enumerate by hand. hl.tile([m, n]) tells Helion to tile that iteration space; it does not say how big the tiles are, in what order they run, or how memory is addressed. Those become autotuner decisions.

The first time you call an un-configured Helion kernel, it runs a real search. Meta's own example output from that process:

[586s] Autotuning complete in 586.6s after searching 1520 configs.
One can hardcode the best config and skip autotuning with:
    @helion.kernel(config=helion.Config(block_sizes=[64, 64, 64],
loop_orders=[[0, 1]], l2_groupings=[4], range_unroll_factors=[0, 1],
range_warp_specializes=[None, False], range_num_stages=[0, 3],
range_multi_buffers=[None, False], range_flattens=[None, None],
num_warps=8, num_stages=6, indexing='block_ptr', pid_type='flat'))

That's nearly 10 minutes spent evaluating 1,520 candidate Triton configurations for one kernel shape, which is exactly the kind of GPU-bound, bursty workload that per-minute billing with no minimum commitment is built for: you're paying for a search, not a steady-state job. Once the search finishes, you paste the printed config into the decorator and every subsequent run skips the search entirely, compiling straight to that one configuration.

The table below is the actual search space Helion's autotuner explores behind that single hl.tile call, reconstructed from Meta's own configuration reference:

ParameterWhat it controls
indexingMemory access strategy: pointer arithmetic, block pointers, or tensor descriptors (which use Tensor Memory Accelerators on Hopper/Blackwell)
block_sizesTile size per dimension in an hl.tile loop, affecting register usage and parallelism
flatten_loopsWhether to flatten a multi-dimensional tile space into one dimension
loop_orders / l2_groupingIteration order of nested tiles and PID swizzling for L2 cache reuse
reduction_loopsPersistent (single-pass) versus looped reduction strategy
pid_typeGrid mapping strategy: flat 1D, multi-dimensional, or persistent kernels pinned one block per SM
load_eviction_policyL1 cache residency hints on Triton loads
num_warps, num_stages, and related range knobsStandard Triton tunables the autotuner sweeps directly

That's the search space this post's angle is really about: a single hl.tile call is shorthand for "try all of these," where a hand-written Triton kernel would need that entire table enumerated, one axis at a time, inside a triton.autotune decorator, and rewritten whenever a new indexing strategy needs testing.

Is Helion Actually Faster to Write Than Triton? What the Line Counts Show

The productivity claim has a concrete number behind it, and it's not a rounding trick: Meta states that "the Attention kernel is just 30 lines in Helion, compared to 120 lines in Triton and thousands of lines in CUDA." That's exactly 4x fewer lines than hand-written Triton for the same kernel, and two orders of magnitude fewer than CUDA. The reduction isn't from Helion doing less, it's from the autotuner absorbing everything in that configuration table above: instead of writing three separate code paths for pointer arithmetic, block pointers, and tensor descriptors, you write one and let the search try all three.

Meta's writeup doesn't print the attention kernel's own code, so that 4x figure can't be checked against the source it describes. The matmul kernel earlier in this post can be checked, because the full Helion source is already on the page. Here's a functionally equivalent hand-written Triton kernel for the same operation, written for this post rather than copied from a tutorial, covering the same ground Helion's hl.tile calls absorb automatically: stride-based pointer arithmetic, an explicit autotune config list, and manual masking on every load and the final store.

python
import torch
import triton
import triton.language as tl

@triton.autotune(
    configs=[
        triton.Config({"BLOCK_M": 64, "BLOCK_N": 64, "BLOCK_K": 32}, num_warps=4, num_stages=3),
        triton.Config({"BLOCK_M": 128, "BLOCK_N": 64, "BLOCK_K": 32}, num_warps=8, num_stages=3),
        triton.Config({"BLOCK_M": 64, "BLOCK_N": 128, "BLOCK_K": 64}, num_warps=4, num_stages=4),
    ],
    key=["M", "N", "K"],
)
@triton.jit
def matmul_kernel(
    x_ptr, y_ptr, out_ptr,
    M, N, K,
    stride_xm, stride_xk,
    stride_yk, stride_yn,
    stride_om, stride_on,
    BLOCK_M: tl.constexpr, BLOCK_N: tl.constexpr, BLOCK_K: tl.constexpr,
):
    pid_m = tl.program_id(0)
    pid_n = tl.program_id(1)
    offs_m = pid_m * BLOCK_M + tl.arange(0, BLOCK_M)
    offs_n = pid_n * BLOCK_N + tl.arange(0, BLOCK_N)
    offs_k = tl.arange(0, BLOCK_K)
    x_ptrs = x_ptr + offs_m[:, None] * stride_xm + offs_k[None, :] * stride_xk
    y_ptrs = y_ptr + offs_k[:, None] * stride_yk + offs_n[None, :] * stride_yn
    acc = tl.zeros((BLOCK_M, BLOCK_N), dtype=tl.float32)
    for k in range(0, K, BLOCK_K):
        x_mask = (offs_m[:, None] < M) & (offs_k[None, :] + k < K)
        y_mask = (offs_k[:, None] + k < K) & (offs_n[None, :] < N)
        x = tl.load(x_ptrs, mask=x_mask, other=0.0)
        y = tl.load(y_ptrs, mask=y_mask, other=0.0)
        acc = tl.dot(x, y, acc)
        x_ptrs += BLOCK_K * stride_xk
        y_ptrs += BLOCK_K * stride_yk
    out_ptrs = out_ptr + offs_m[:, None] * stride_om + offs_n[None, :] * stride_on
    out_mask = (offs_m[:, None] < M) & (offs_n[None, :] < N)
    tl.store(out_ptrs, acc, mask=out_mask)


def matmul(x: torch.Tensor, y: torch.Tensor) -> torch.Tensor:
    M, K = x.shape
    K, N = y.shape
    out = torch.empty((M, N), device=x.device, dtype=x.dtype)
    grid = lambda meta: (triton.cdiv(M, meta["BLOCK_M"]), triton.cdiv(N, meta["BLOCK_N"]))
    matmul_kernel[grid](
        x, y, out,
        M, N, K,
        x.stride(0), x.stride(1),
        y.stride(0), y.stride(1),
        out.stride(0), out.stride(1),
    )
    return out

Counted the same way for both, non-blank lines only: the Helion matmul above runs 12 lines of actual logic (14 counting its two comments). This Triton equivalent needs 52, almost entirely spent on the three things hl.tile replaced: the autotune config list, the stride arithmetic for every pointer, and the masking on each load and the store. That's roughly a 4x gap on a second kernel, independently counted, which lines up with Meta's attention-kernel ratio instead of being an artifact of which kernel got chosen for the blog post.

What this check doesn't cover is runtime. Confirming the two kernels compile to comparable performance means running both on the same hardware with the same input shapes, a GPU-bound benchmark this post didn't execute. The performance figures below remain Meta's own published numbers, labeled as such throughout; the line-count side of the productivity claim is the part an independent count on a second kernel can actually confirm without a GPU in hand.

Line count is a proxy for maintenance burden, not runtime speed, so the more important question is whether the generated kernel actually performs. Meta's published benchmarks compare Helion, torch.compile with max-autotune, and hand-written Triton against eager-mode PyTorch across a wide set of kernels, on both NVIDIA and AMD hardware:

GPUHelion geomeantorch.compile geomeanHand-written Triton geomean
NVIDIA B2003.27x2.7x1.76x
AMD MI350X2.37x2.26x1.65x

On B200, that's Helion beating hand-written Triton by 1.85x on average and torch.compile by 1.21x; on MI350X, 1.44x over Triton and 1.05x over torch.compile. The spread across individual kernels is wide: Helion's softmax kernel beat torch.compile by 2.28x on B200, its jsd kernel beat hand-written Triton by 6.22x on the same GPU, and on MI350X the int4_gemm and jsd kernels beat Triton by 4.5x and 4.4x respectively. Those are the outliers, not the median, so treat any single kernel's multiplier as a best case rather than a baseline you should expect by default.

There's a second data point worth taking seriously: a Helion implementation of an RMSNorm backward kernel, described in the same writeup as "written in less than a day," matched or exceeded a hand-optimized CuTe DSL kernel from the Quack library on H100 across a range of reduction dimensions. That's the productivity argument in its sharpest form: expert-level performance without the multi-day tuning cycle CuTe DSL typically demands. It's also, worth noting, a comparison Meta ran and published itself, not an independent reproduction, so read the specific multipliers as the vendor's own best-case numbers rather than a neutral third-party audit.

Choosing a Kernel Stack for a Multi-GPU-Vendor Fleet (NVIDIA, AMD, Intel)

Here's the question this post set out to answer: if you're running kernels across more than one GPU vendor, can you actually standardize on a single Helion source file, or is that still marketing?

The honest answer is: mostly yes, with real caveats. Because Helion compiles to Triton rather than replacing it, it inherits Triton's own backend coverage, and our Triton on AMD ROCm vs NVIDIA CUDA benchmark already shows that the same Triton kernel source runs unmodified on both platforms, with autotuning as the part that doesn't transfer. Helion doesn't remove that autotuning gap, it automates it: instead of manually re-running a triton.autotune sweep with vendor-specific search spaces when you move from H100 to MI300X, Helion's autotuner just runs its search again on whichever GPU you're currently on and produces a new config for that hardware.

Triton itself already has a third vendor path: intel-xpu-backend-for-triton, a community-maintained project describing itself as an "OpenAI Triton backend for Intel GPUs." Because a Helion kernel's device code compiles to Triton, that backend is architecturally reachable from Helion, but it isn't a path Helion's own repository documents or tests.

What this means for picking hardware to validate on:

  • NVIDIA and AMD are both live, tested paths today. You can write one Helion kernel, run it on an H100 rental, let the autotuner search, then run the same source on an MI300X rental and let it search again for that hardware. Spheron's own catalog covers exactly this pair: $2.65/hr on-demand for H100 and $3.59/hr on-demand for MI300X (as of 17 Sep 2026), both provisioning in under 2 minutes with per-minute billing after a 20-minute minimum. For a kernel whose first autotuning run alone can take around 10 minutes, that billing model matters: you're not paying for a monthly reservation to run a search you do once and then cache.
  • Intel is the honest gap. Intel engineers are credited among Helion's own contributors, and Triton has its own separate Intel GPU backend, but Helion's repository doesn't document a tested Intel path the way it documents its NVIDIA extras, so it's earlier-stage than the CUDA and ROCm routes. It also isn't something you can validate on Spheron: Spheron's marketplace lists NVIDIA and AMD MI300X GPUs, not Intel hardware. If the full three-vendor portability story is what you're testing, the Intel leg needs a different provider.
  • Spot pricing changes the calculus for kernel iteration, not just production serving. Autotuning a kernel is bursty, GPU-bound work with no state to lose between runs, which is exactly the profile spot capacity fits, since on-demand and spot rates move with availability day to day.

For actually spinning up either GPU family to try this, Spheron's documentation is the starting point for account setup and instance access.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 17 Sep 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

The realistic verdict: Helion's multi-backend promise holds for the two vendors that matter most in production GPU rental today, NVIDIA and AMD, because it's standing on Triton's existing portability rather than inventing a new one. The three-vendor story is real in who's credited as a contributor and in Triton's own separate Intel backend, but Helion itself doesn't yet document or test an Intel path, making it the newest, least-proven leg of the story.

If you're validating a Helion kernel's autotuned config across GPU vendors, Spheron rents both NVIDIA H100 (on-demand and spot) and AMD MI300X (on-demand), with per-minute billing after a 20-minute minimum and provisioning in under 2 minutes, no contract needed to run a 10-minute autotuning search and walk away.

Rent H100 or MI300X on Spheron →

FAQ / 05

Frequently Asked Questions

A PyTorch-native kernel compiler takes code written in PyTorch's own tensor syntax, rather than a separate low-level language, and lowers it into a fast GPU kernel. torch.compile and its TorchInductor backend do this automatically for existing PyTorch models. Helion extends the same idea to hand-authored kernels: you write tile-level logic with ordinary PyTorch operators and hl.tile loops, and Helion's compiler, which reuses TorchInductor for the final lowering, turns it into autotuned Triton code.

No, but they share machinery. torch.compile automatically captures and optimizes an existing PyTorch model's graph with no code changes required. Helion is for kernels you write deliberately, using explicit hl.tile loops to mark the parallel iteration space. Helion's own compiler pipeline calls TorchInductor to map the PyTorch operators inside a kernel body to Triton, so TorchInductor is a component Helion depends on, not a competing tool.

No. Helion compiles down to Triton rather than replacing it, so every Helion kernel is a Triton kernel underneath, generated automatically instead of hand-written. Triton stays the right choice when you need to hand-place a specific memory access pattern the autotuner doesn't already search, or when you're maintaining an existing Triton codebase. Helion is the right choice when you want the autotuner to search block sizes, loop order, and memory layout for you.

Helion compiles to Triton, so it inherits Triton's backend coverage: NVIDIA CUDA and AMD ROCm are both production paths today. Triton also has a separate, community-maintained project, intel-xpu-backend-for-triton, describing itself as an "OpenAI Triton backend for Intel GPUs," though Helion's own repository doesn't document an Intel path.

In Meta's own published benchmarks, yes on average, though not on every kernel. Across a wide set of kernels on NVIDIA B200, Helion's geomean speedup over eager-mode execution was 3.27x against 1.76x for hand-written Triton, a 1.85x edge. On AMD MI350X, Helion's geomean was 2.37x against Triton's 1.65x, a 1.44x edge. Individual kernels varied widely: Helion beat hand-written Triton by 6.22x on a jsd kernel on B200, and by 4.5x on an int4_gemm kernel on MI350X.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min