Skip to content

Faster AVX2 swizzle_dyn(_precise) - #362

Merged
Shnatsel merged 2 commits into
linebender:mainfrom
Shnatsel:faster-avx2-swizzle
Aug 31, 2026
Merged

Faster AVX2 swizzle_dyn(_precise)#362
Shnatsel merged 2 commits into
linebender:mainfrom
Shnatsel:faster-avx2-swizzle

Conversation

@Shnatsel

Copy link
Copy Markdown
Contributor

@dzaima got nerd-sniped by my blog post about swizzle_dyn_precise and workshopped even faster formulations. I couldn't let that stand and tried optimizing these this myself.

All the formulations we tried, their godbolt links and their llvm-mca timings can be found at https://gist.github.com/Shnatsel/38abb51f0dc337837b1c84191461ed56

The one proposed in this PR is the last line of the table, "Blend using control-derived mask".

On my Zen2 laptop this improves performance dramatically: 33% less time taken which means 50% higher throughput.

Corresponding std::simd PR: rust-lang/portable-simd#548


This PR also includes an optimization to swizzle_dyn, the non-precise variant. Details on its performance are in the commit message.

40% faster on Haswell
unchanged on Skylake
about 26% faster on Alder Lake P
25% faster on Zen 1 and Zen 3

Uses one more register (up from 4 to 5 peak YMM registers) for an added constant

Gets folded into a `shufflevector` for constant indices either way; it reaches x86 backend as two shuffles with OR instead of a shufflevector, but gets folded into a shufflevector inside the x86 backend anyway, so the optimizations for constant indices are unaffected.
@Shnatsel
Shnatsel added this pull request to the merge queue Aug 31, 2026
Merged via the queue into linebender:main with commit 7d5835c Aug 31, 2026
22 checks passed
@Shnatsel
Shnatsel deleted the faster-avx2-swizzle branch August 31, 2026 18:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants