Faster AVX2 swizzle_dyn(_precise) - #362
Merged
Merged
Conversation
40% faster on Haswell unchanged on Skylake about 26% faster on Alder Lake P 25% faster on Zen 1 and Zen 3 Uses one more register (up from 4 to 5 peak YMM registers) for an added constant Gets folded into a `shufflevector` for constant indices either way; it reaches x86 backend as two shuffles with OR instead of a shufflevector, but gets folded into a shufflevector inside the x86 backend anyway, so the optimizations for constant indices are unaffected.
LaurenzV
approved these changes
Aug 31, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
@dzaima got nerd-sniped by my blog post about
swizzle_dyn_preciseand workshopped even faster formulations. I couldn't let that stand and tried optimizing these this myself.All the formulations we tried, their godbolt links and their llvm-mca timings can be found at https://gist.github.com/Shnatsel/38abb51f0dc337837b1c84191461ed56
The one proposed in this PR is the last line of the table, "Blend using control-derived mask".
On my Zen2 laptop this improves performance dramatically: 33% less time taken which means 50% higher throughput.
Corresponding
std::simdPR: rust-lang/portable-simd#548This PR also includes an optimization to swizzle_dyn, the non-precise variant. Details on its performance are in the commit message.