Skip to main content

RING_KERNEL_PTX_TEMPLATE

Constant RING_KERNEL_PTX_TEMPLATE 

Source
pub const RING_KERNEL_PTX_TEMPLATE: &str = r#"
.version 8.0
.target sm_75
.address_size 64

.visible .entry ring_kernel_main(
    .param .u64 control_block_ptr,
    .param .u64 input_queue_ptr,
    .param .u64 output_queue_ptr,
    .param .u64 shared_state_ptr
) {
    .reg .u64 %cb_ptr;
    .reg .u32 %one;

    // Load control block pointer
    ld.param.u64 %cb_ptr, [control_block_ptr];

    // Mark as terminated immediately (offset 8)
    mov.u32 %one, 1;
    st.global.u32 [%cb_ptr + 8], %one;

    ret;
}
"#;
👎Deprecated since 1.1.1:

hardcodes sm_75, breaking on Pascal and older GPUs — use ring_kernel_ptx_template_for(major, minor) with the real device’s compute capability instead

Expand description

PTX kernel source template for persistent ring kernel — deprecated fixed-sm_75 fallback, kept only for API compatibility with existing callers that don’t have a device handle available.

PTX built for sm_75 will fail to load (CUDA_ERROR_INVALID_PTX) on any GPU with a compute capability below 7.5 (e.g. Pascal, sm_61) — PTX forward-compatibility only extends to equal-or-newer architectures. Prefer ring_kernel_ptx_template_for with the actual target device’s compute capability (CudaDevice::compute_capability, not part of this crate’s public API surface).