Skip to content

Store GPU particles as struct-of-arrays, use opt.backend in solve loops - #143

Closed
AdityaPandeyCN wants to merge 2 commits into
SciML:mainfrom
AdityaPandeyCN:soa-particles
Closed

AdityaPandeyCN wants to merge 2 commits into
SciML:mainfrom
AdityaPandeyCN:soa-particles

Conversation

@AdityaPandeyCN

Copy link
Copy Markdown
Contributor

Checklist

  • Appropriate tests were added
  • Any code changes were done in a way that does not break public API
  • All documentation related to code changes were updated
  • The new code follows the
    contributor guidelines, in particular the SciML Style Guide and
    COLPRAC.
  • Any new documentation only uses public API

Additional context

Two changes to the GPU swarm solvers (ParallelSyncPSOKernel, ParallelPSOKernel, and the swarm phase of HybridPSO).

1. Struct-of-arrays particle storage on GPU backends. Particles are now kept in one n × (3D + 2) matrix instead of an array of SPSOParticle. The new SoAParticles type is an AbstractVector{SPSOParticle}, so the kernels still do particles[i] and particles[i] = p and are unchanged. CPU backends keep the old storage.

2. vectorized_solve! uses opt.backend. The GPU methods used get_backend(gpu_particles), which builds a fresh default backend and drops whatever options the user passed, e.g. CUDABackend(always_inline = true). The CPU-only method is left alone.

Why

I profiled with Nsight Compute on a T4 (BBOB f8, D = 10, 50k particles):

  • With a 256-byte struct per particle, a warp reading one field hit ~31 sectors per request instead of 8. With this PR it's 7.7.
  • Each thread was also copying the whole prob (~1.5 KB, mostly the objective closure) to its stack because of a non-inlined call in the objective chain. always_inline = true removes it (update kernel stack 1536 B → 32 B), but it only reached the update kernel after change 2.

Results

T4, BBOB at D = 10, 50k particles, ms per iteration:

main this PR this PR + always_inline
SyncPSOKernel, f1 1.08 0.89 0.49
SyncPSOKernel, f8 1.15 0.76 0.57
SyncPSOKernel, f15 5.08 3.82 3.56
PSOKernel, f1 1.02 0.77 0.38
PSOKernel, f8 0.88 0.84 0.46
PSOKernel, f15 3.61 3.38 3.34

…ve loops

- src/soa.jl: SoAParticles (one n x (3D+2) matrix) for GPU backends, with
  particle_storage, minimum and get_positions; CPU storage unchanged.
- init/solve for ParallelSyncPSOKernel and ParallelPSOKernel use it.
- GPU vectorized_solve! methods use opt.backend instead of rebuilding a
  default backend, so options like always_inline are kept.
- README: recommend CUDABackend(always_inline = true).
- test: CPU round-trip test for SoAParticles.
@AdityaPandeyCN

Copy link
Copy Markdown
Contributor Author

@utkarsh530 Please take a look

@AdityaPandeyCN

Copy link
Copy Markdown
Contributor Author

Closing this, will attempt in a better way.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant