Hacker News new | ask | show | jobs
by thrtythreeforty 21 days ago
Surely SIMD combined with multiple streams would beat both approaches. (This would be separate streams in each SIMD lane and separate streams in different SIMD variables.) There are multiple SIMD execution units, just like the 6 scalar units you mention. The latency of SIMD ops will be similar to scalar, except in cases you mention like shifts.
1 comments

gonna try it! i also suspect that that cranelift's vector lowering is not ideal, so the reasons i described could be wrong