NVIDIA Introduces CUDA Rust for Writing GPU Kernels

NVIDIA Introduces CUDA Rust for Writing GPU Kernels

NVIDIA is adding native Rust support for writing GPU kernels, through two new projects: cuda-oxide and cutile-rs.

CUDA has been effectively C++ only since it appeared. Bindings from other languages have always existed, and they let you call GPU code rather than write it. The kernel itself, the function that runs on the GPU, has been C++.

Why this is more than a language preference

GPU kernels are an unusually dangerous place to write C++.

There is no memory protection on the device in the sense CPUs have. A kernel indexing past the end of a buffer does not segfault; it reads or writes whatever is adjacent in GPU memory. The result is silently wrong numbers, and you find out much later, if at all.

Debugging is correspondingly miserable. The usual tools apply awkwardly, thousands of threads are executing concurrently, and the failure may not be deterministic.

The bug classes Rust’s type system eliminates at compile time, out-of-bounds indexing, use-after-free, data races, are exactly the ones that are hardest to find on a GPU. That makes this a better fit than “Rust everywhere” enthusiasm would suggest on its own.

Whether it delivers depends entirely on how much of Rust’s safety survives contact with GPU execution. Shared memory between threads in a block, warp-level primitives, and manual synchronisation are all things the borrow checker was not designed around, and any of them may need unsafe. A Rust GPU story where the interesting half is unsafe blocks is worth considerably less than the announcement implies. That is the question to ask when the projects mature.

The Rust-in-systems pattern

This lands alongside a broader move that has been running for a while.

uutils coreutils is the default on Ubuntu. sudo-rs exists. Canonical is funding C-to-Rust translation research. Rust for Linux continues upstream.

Those are all rewrites of existing C. This is different: a new development surface where the vendor is offering Rust alongside the incumbent rather than replacing it. Nobody has to port anything, which removes the usual objection that rewriting mature tested code trades known bugs for unknown ones.

It also comes from NVIDIA rather than from the community, which matters for adoption. A language binding maintained by volunteers against a proprietary toolchain is a support risk; one from the vendor is a different proposition.

Tempering it

These are new projects. “Introduces” is not “production ready,” and the CUDA ecosystem, cuDNN, cuBLAS, TensorRT, the profilers, and twenty years of accumulated example code, is all C++. A Rust kernel that has to interoperate with that is doing translation at every boundary.

The realistic near-term audience is people writing custom kernels from scratch for numerical or graphics work, not people maintaining an existing CUDA codebase.

For anyone renting GPUs to run this on, our GPU hosting providers directory covers the options, and our Ollama guide covers the inference side, which is the layer most people interact with rather than writing kernels themselves.