Expand description
Portable, auto-vectorized SHA-256 multi-lane kernel.
A round reads the state the round before it wrote, so one chain stalls on its own latency. Every entry point here advances several independent chains at once instead.
The lanes are held transposed: each state and message word becomes one word per lane, and every step is a fixed-width loop over the lanes.
No intrinsics and no unsafe code, so the vectorizer fills whatever width the target has. A hand-written kernel takes over only at the lane counts it claims.
These loops are also the reference the kernels are tested against.
Reference: FIPS 180-4, section 6.2.
Constants§
- LANES
- Batch width the dispatch is tuned for on this target.
Functions§
- compress256_
multi - Compresses one 64-byte block into each of a batch of independent SHA-256 states, in place.
- compress256_
multi_ portable - Compresses one 64-byte block into each state of a batch, with plain lane loops.
- sha256_
multi - Hashes a batch of equal-length byte inputs into one standard SHA-256 digest each.