the early versions of L (circa ... 2011/2012) couldn't really deliver any substantial performance improvement over what is out there (BQN/ngn/Kona). The memory bandwidth was the limit - I simply couldn't keep the cores active. ~2018/2019 I started from scratch with arrays (vectors) using compression by default. That was a massive unlock - and I could not keep most of the CPU doing actual compute! Then it was years of working on compression native operations - some of which were obvious[3] like sum/reductions ... others not so much!
the early versions of L (circa ... 2011/2012) couldn't really deliver any substantial performance improvement over what is out there (BQN/ngn/Kona). The memory bandwidth was the limit - I simply couldn't keep the cores active. ~2018/2019 I started from scratch with arrays (vectors) using compression by default. That was a massive unlock - and I could not keep most of the CPU doing actual compute! Then it was years of working on compression native operations - some of which were obvious[3] like sum/reductions ... others not so much!
[1] https://www.emergentmind.com/topics/memory-wall [2] https://www.cse.iitd.ac.in/~rijurekha/col216/quantitative_ap... [3] https://lv1.sh/blog/compute-on-compressed/