Skip to content

Remove kernel keyword arguments that caused a per-thread copy of prob - #145

Merged
ChrisRackauckas merged 1 commit into
SciML:mainfrom
AdityaPandeyCN:kernel-stack-copy
Sep 30, 2026
Merged

ChrisRackauckas merged 1 commit into
SciML:mainfrom
AdityaPandeyCN:kernel-stack-copy

Conversation

@AdityaPandeyCN

Copy link
Copy Markdown
Contributor

The update kernels had c1/c2 keyword arguments that no caller passes. Keyword arguments make Julia split the kernel into a wrapper and a separate body function, and on GPU that body function was not inlined, so every thread copied prob onto its stack before calling it.

This replaces them with constants (PSO_C1, PSO_C2). It also adds a small inlined rand_static for the SVector random draws, since StaticArrays' rand was not inlined either and used stack for its return value.

Stack per thread from ptxas -v (sm_75, D = 10, default CUDABackend()):

Objective Kernel main this PR main + always_inline
Rosenbrock ParallelSyncPSOKernel 336 B 0 B 0 B
Rosenbrock ParallelPSOKernel 352 B 32 B 32 B
BBOB f8 ParallelSyncPSOKernel 1216 B 1008 B 0 B

For BBOB the remaining stack comes from the objective itself (BBOBFunction holds ~844 B of data and its call is not inlined), so that part still needs always_inline = true.

CPU Core tests pass locally. I haven't run the CUDA tests or timed it on a GPU.

The c1/c2 keyword arguments on the update kernels were never passed by any caller, but they made Julia split each kernel into a wrapper and a separate body function. That body function was not inlined on GPU, so every thread copied prob onto its stack. Use constants instead, and add an inlined rand_static for the SVector random draws.
@ChrisRackauckas
ChrisRackauckas merged commit f47a0b6 into SciML:main Sep 30, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants