It is a fun way to benchmark your system. I managed to get 4.1hz on my 13th gen i5 work laptop.
Let's see an M-series chip blow that out of the water. ;)
EDIT - It was 4.1Hz in advance mode. Actually runs at between 22Hz and 44hz depending on each time I run it. Odd to have such a wide performance variation. I will just blame Windows and some power management setting/thermal throttling.
M4 Max, Brave — 93Hz highest, usually around 90Hz. On Advanced it started at 22Hz & steadily dropped from there into the single digits as it ran. Safari held around 62.xHz. On Advanced it started at 40Hz & also kept dropping.
So, yeah, this is wonderful since the ARM architecture (Acorn Risc Machine) is arguably derived from the 6502 by Sophie Wilkins and Steve Furber at Acorn Computers.
Shameless plug: I wrote a 'remix' a couple of years ago over a Christmas break using WASM+WebGL+DearImGui with better rendering performance, integrated assembler and a couple visualization tools (using the visual6502 team's original reverse engineered netlist data, and https://github.com/mist64/perfect6502 as simulation loop):
The link to Perfect6502 is interesting. It says that it supposedly gets speeds of 1/30th of 1MHz - i.e. ~30KHz - on high-end 2025 CPUs. Which is very slow of course, but still ~1000x faster than what people are reporting in this thread.
Of course it's in JS rather than C, but V8 is really good, well within an order of magnitude - not three - of native, at least with well-written code. I profiled the original Visual6502 and at a glance it sure does look like it's losing a ton of performance to rendering (so good call there).
I have to wonder what the limit is. Even Perfect6502 looks like it's all single-threaded and CPU-driven, which really isn't making the most of modern machines. But I don't know if this is a task that could be efficiently split across cores or dispatched to the GPU (and I can't even begin to know whether SIMD might be useful); of course hardware itself is fundamentally parallel, but the synchronization might kill you?
As far as I understood it, this basically re-evaluates the state of a 'dirty list' of transistors though all their connections, and loops until the node list has 'settled' again into a stable state (e.g. state changes from the last loop didn't cause trigger new state changes). Haven't thought about how this could be threaded or simd-ified.
Similar project, gate level emulation of NES CPU+PPU by the Nesticle author Icer Addis.
An entire world in a scanline