Repository navigation
Store GPU particles as struct-of-arrays, use opt.backend in solve loops - #143
Closed
AdityaPandeyCN wants to merge 2 commits into
Closed
AdityaPandeyCN wants to merge 2 commits into
AdityaPandeyCN wants to merge 2 commits into
Conversation
…ve loops - src/soa.jl: SoAParticles (one n x (3D+2) matrix) for GPU backends, with particle_storage, minimum and get_positions; CPU storage unchanged. - init/solve for ParallelSyncPSOKernel and ParallelPSOKernel use it. - GPU vectorized_solve! methods use opt.backend instead of rebuilding a default backend, so options like always_inline are kept. - README: recommend CUDABackend(always_inline = true). - test: CPU round-trip test for SoAParticles.
Contributor
Author
|
@utkarsh530 Please take a look |
Contributor
Author
|
Closing this, will attempt in a better way. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Checklist
contributor guidelines, in particular the SciML Style Guide and
COLPRAC.
Additional context
Two changes to the GPU swarm solvers (ParallelSyncPSOKernel, ParallelPSOKernel, and the swarm phase of HybridPSO).
1. Struct-of-arrays particle storage on GPU backends. Particles are now kept in one
n × (3D + 2)matrix instead of an array ofSPSOParticle. The newSoAParticlestype is anAbstractVector{SPSOParticle}, so the kernels still doparticles[i]andparticles[i] = pand are unchanged. CPU backends keep the old storage.2.
vectorized_solve!usesopt.backend. The GPU methods usedget_backend(gpu_particles), which builds a fresh default backend and drops whatever options the user passed, e.g.CUDABackend(always_inline = true). The CPU-only method is left alone.Why
I profiled with Nsight Compute on a T4 (BBOB f8, D = 10, 50k particles):
prob(~1.5 KB, mostly the objective closure) to its stack because of a non-inlined call in the objective chain.always_inline = trueremoves it (update kernel stack 1536 B → 32 B), but it only reached the update kernel after change 2.Results
T4, BBOB at D = 10, 50k particles, ms per iteration:
always_inline