Skip to main content

Crate ringkernel_cuda

Crate ringkernel_cuda 

Source
Expand description

CUDA Backend for RingKernel

This crate provides NVIDIA CUDA GPU support for RingKernel using cudarc.

§Features

  • Persistent kernel execution (cooperative groups)
  • Lock-free message queues in GPU global memory
  • PTX compilation via NVRTC
  • Multi-GPU support

§Requirements

  • NVIDIA GPU with Compute Capability 6.0+ (Pascal and newer) for the core persistent-actor/cooperative-groups path. Some features have higher floors (e.g. Thread Block Clusters/DSMEM/TMA/Green Contexts require Hopper, CC 9.0+) — see the top-level README’s “Feature to minimum compute capability” table for the full breakdown.
  • CUDA Toolkit 11.0+
  • Native Linux (persistent kernels) or WSL2 (event-driven fallback)

§Example

ⓘ
use ringkernel_cuda::CudaRuntime;
use ringkernel_core::runtime::RingKernelRuntime;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let runtime = CudaRuntime::new().await?;
    let kernel = runtime.launch("vector_add", Default::default()).await?;
    kernel.activate().await?;
    Ok(())
}

Modules§

stub 🔒

Structs§

CudaRuntime
Stub runtime when the backend feature is disabled.

Constants§

RING_KERNEL_PTX_TEMPLATEDeprecated
PTX kernel source template for persistent ring kernel — deprecated fixed-sm_75 fallback, kept only for API compatibility with existing callers that don’t have a device handle available.

Functions§

compile_ptx
Stub compile_ptx when CUDA is not available.
cuda_device_count
Get CUDA device count.
is_cuda_available
Check if CUDA is available at runtime.
ring_kernel_ptx_template_for
PTX kernel source template for persistent ring kernel, generic over the target compute capability (major, minor).