Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Isn't there something to be said for owning your own hardware though?


Not if it's 5-10x slower than a remote inference server. Mac prefill latency is exhausting.


Currently m5 max has a prefill rate for Qwen 3.8 27B of 400+ tok/s, then generate at ~60 tok/s. (And it can be improved further, the software is not yet at the level it is for CUDA)

That means for context under ~5k or so ttft (time to first token) it's going to respond faster than Claude. If the answer is less than ~1k I think the request finishes sooner. And it's ~claude 4.5 or 4.6 level intelligence.

I've had it work for more than a day on a pi.dev "loop engineering" project involving writing software.

Plus privacy. Plus offline.


M5 changed that a lot though - it could still be better but 4x improvement made it cross the frustratingly slow barrier for me.


Oh tell me more about prefill latency.



Thank you very much for this link! Extremely fascinating.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: