Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Yes, but “CUDA for Rust” currently covers two different jobs, and the difference determines which tool you need. Rust code running on the CPU can call CUDA’s APIs to select a GPU, manage memory, and launch work; the cudarc crate is a Rust binding for that layer. The newer development is writing the GPU kernel itself in Rust. NVIDIA announced two native kernel tracks on September 8, 2026: cuda-oxide, which uses a SIMT (single instruction, multiple threads) model, and cuTile Rust, which uses a tile-based model. Both are real, but they are at different maturity levels, and their requirements differ.

Two layers sit under the phrase “CUDA for Rust”

CUDA is NVIDIA’s GPU programming platform. The CUDA Programming Guide defines it as “a parallel computing platform and programming model developed by NVIDIA that enables dramatic increases in computing performance by harnessing the power of the GPU.” Because it is tied to NVIDIA hardware, every project in this guide needs an NVIDIA GPU, and none of them run on other vendors’ chips.

For a Rust developer, the work splits into two layers:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Host code runs on the CPU and controls the GPU. It selects a device, allocates memory, transfers data, and launches kernels. Bindings such as cudarc expose CUDA’s APIs to Rust for this job.
  • Device code is the kernel, which runs on the GPU. Rust developers could already launch kernels from Rust, but they often had to write the kernel itself in another language. The newer projects target that gap by compiling Rust kernel code for the GPU.

Most questions about “CUDA in Rust” come down to which of these layers you need to write.

#1 Best Overall
Sale
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
  • AI Performance: 767 AI TOPS
  • OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis

Which project covers which need

Project Layer Kernel model Maturity, as described in its own material
cudarc Host-side Rust bindings for CUDA APIs Does not author kernels Not stated
cuda-oxide Device kernels written in Rust SIMT, compiled to PTX through a custom backend The project book labels v0.1.0 an “early-stage alpha”
cuTile Rust Device kernels written in Rust Tile-based, mapped through CUDA Tile IR NVIDIA describes the larger effort as continuing to mature and says development continues into 2027 and beyond
Rust-CUDA Rust-to-PTX compilation plus CUDA ecosystem libraries SIMT-style kernels compiled to PTX The project guide describes an effort to make Rust a tier-1 language for GPU computing; release status not stated
CubeCL Cross-vendor GPU compute, per NVIDIA’s ecosystem appendix Portability and DSL-oriented goals Not CUDA-specific; evaluate separately

SIMT means you write code for individual threads, and the hardware executes them in groups. Tile-based programming means you describe operations on blocks of data instead.

cuda-oxide: SIMT kernels in Rust

cuda-oxide compiles Rust kernel code to PTX through a custom backend. Its model is the one CUDA C++ programmers already use, so it suits teams whose existing kernels are thread-oriented and who want Rust’s language features in device code.

Maturity is the main caveat. The cuda-oxide book describes v0.1.0 as “an early-stage alpha” and warns that it may contain bugs, incomplete features, and API breakage. Expect to revise code as the project changes.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cuTile Rust: tile-based kernels

cuTile Rust takes a different approach. Instead of writing per-thread logic, you express work as operations on tiles, which are blocks of data, and the tile model is mapped through CUDA Tile IR. This fits workloads that are naturally blocked array operations, such as matrix multiplication.

The most specific performance figures available come from the 2026 paper Fearless Concurrency on the GPU. For cuTile Rust on an NVIDIA B200, the authors report 7 TB/s for element-wise operations and 2 PFlop/s for GEMM, which they put at 96% of cuBLAS. These numbers come from that device and those workloads. They say nothing about other GPUs, other kernels, or other problem sizes, and they are not a guarantee for your code.

Earlier and complementary projects

Rust-CUDA

The Rust-CUDA project guide describes an effort to make Rust a tier-1 language for GPU computing with CUDA. That includes tooling to compile Rust to PTX and to use CUDA libraries from Rust. Its setup page notes that the LLVM 7.x requirement can make installation difficult and points to Docker images that include CUDA and LLVM, which is a practical way around local toolchain problems.

cudarc

cudarc provides Rust bindings to CUDA APIs. It is the right choice when the host program is the part you are writing in Rust and you need to drive the GPU from it. It does not let you write kernel code in Rust, so pair it with a kernel from one of the projects above or with kernels that already exist. The guide’s requirements table omits cudarc because the cited material does not give its compatibility baseline; check its own documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4. System Requirements: Minimum 850W PSU with 16-pin 12V-2x6 (12VHPWR) connector required. Verify before purchasing.
  • Military-grade components deliver rock-solid power and longer lifespan for ultimate durability. Compatibility: 348mm (13.7") length, 3.6 slots, 4.3 lbs. Confirm case clearance and slot spacing. GPU bracket included.
  • Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
  • 3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans
  • Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads

Requirements differ by project

Do not treat the following as one “Rust CUDA” minimum. Each column describes one project’s own setup requirements.

Requirement Rust-CUDA (project setup guide) cuda-oxide SIMT track (NVIDIA announcement) cuTile Rust (NVIDIA announcement)
Operating system Not stated Linux Not stated
GPU Compute Capability 5.0 (Maxwell) or later Compute Capability 8.0 or later Not stated
CUDA CUDA 12.0 or newer CUDA Toolkit 12.x or newer Not stated
Driver An appropriate NVIDIA driver Not stated Not stated
Compiler LLVM 7.x clang with libclang headers Not stated
Rust toolchain Not stated A pinned nightly toolchain Not stated

“Not stated” means the project’s cited setup material does not give that requirement, so check its setup page before assuming it is unneeded. A GPU below Compute Capability 8.0 can meet Rust-CUDA’s stated baseline and still fall outside the cuda-oxide track.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Installing the CUDA Toolkit on Linux

NVIDIA’s installation guide documents three Linux routes for the CUDA Toolkit:

  • Package manager installation
  • Runfile installer
  • Conda

The guide also covers pip wheels, which are oriented toward Python runtime use. Check the installation page before relying on them to build kernels.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Version numbers need care. NVIDIA’s CUDA Toolkit documentation landing page highlights CUDA 13.4, but the CUDA Programming Guide it links is Release 13.2. These two may not describe the same release, so take the version from the page you are reading and confirm supported distributions and drivers on NVIDIA’s installation page.

Choosing a path

  1. You only need Rust host code that calls CUDA. Use cudarc for the host side. Your kernels come from elsewhere.
  2. You want kernels written in Rust and your algorithm is thread-oriented. Evaluate cuda-oxide, but only if you run Linux, have a GPU at Compute Capability 8.0 or later, and can maintain a pinned nightly toolchain and absorb API changes in an alpha release.
  3. Your workload is blocked array math, such as GEMM-heavy code. Evaluate cuTile Rust, and benchmark it on your own shapes and hardware rather than relying on the paper’s numbers.
  4. You need Rust-to-PTX compilation with CUDA libraries and can run LLVM 7.x on older hardware. Consider Rust-CUDA, starting from its Docker images if the local LLVM setup fails.
  5. You need one kernel codebase across GPU vendors. Look at portability projects such as CubeCL. CUDA itself is not the path.

Project versions, toolchains, and GPU support change quickly. The most recent source behind this guide is NVIDIA’s September 8, 2026 announcement, so confirm the current state on each project’s own release notes and setup page before you adopt one. For production use, also check issue activity and supported features, and validate the chosen project against your own workload.

Quick Recap

SaleBestseller No. 1
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
AI Performance: 767 AI TOPS; OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode); Powered by the NVIDIA Blackwell architecture and DLSS 4
$786.37
Bestseller No. 2
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans; Auto-Extreme precision automated manufacturing helps ensure higher reliability
$1,831.31

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.