I think pretty much any multicore ARM CPU with a post ARMv8 ISA is looking pretty strong for local AI right now. Same goes for x86 chips with AVX2 support.
Apple Silicon AMX units provide the matrix multiplication performance of many core CPUs or faster at a fraction of the wattage. See eg.
Plus, the benchmark you've linked to is comparing hardware accelerated inferencing to the notoriously crippled MKL execution. A more appropriate comparison would test Apple's AMX units against the Ryzen's AVX-optimized inferencing.
Apple Silicon AMX units provide the matrix multiplication performance of many core CPUs or faster at a fraction of the wattage. See eg.
https://explosion.ai/blog/metal-performance-shaders https://github.com/danieldk/gemm-benchmark#1-to-16-threads