pub const RING_KERNEL_PTX_TEMPLATE: &str = r#"
.version 8.0
.target sm_75
.address_size 64
.visible .entry ring_kernel_main(
.param .u64 control_block_ptr,
.param .u64 input_queue_ptr,
.param .u64 output_queue_ptr,
.param .u64 shared_state_ptr
) {
.reg .u64 %cb_ptr;
.reg .u32 %one;
// Load control block pointer
ld.param.u64 %cb_ptr, [control_block_ptr];
// Mark as terminated immediately (offset 8)
mov.u32 %one, 1;
st.global.u32 [%cb_ptr + 8], %one;
ret;
}
"#;👎Deprecated since 1.1.1:
hardcodes sm_75, breaking on Pascal and older GPUs — use ring_kernel_ptx_template_for(major, minor) with the real device’s compute capability instead
Expand description
PTX kernel source template for persistent ring kernel — deprecated
fixed-sm_75 fallback, kept only for API compatibility with existing
callers that don’t have a device handle available.
PTX built for sm_75 will fail to load (CUDA_ERROR_INVALID_PTX) on any
GPU with a compute capability below 7.5 (e.g. Pascal, sm_61) — PTX
forward-compatibility only extends to equal-or-newer architectures.
Prefer ring_kernel_ptx_template_for with the actual target device’s
compute capability (CudaDevice::compute_capability, not part of this
crate’s public API surface).