Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

gpt-oss-20b: 9 threads, 131072 context window, 4 experts - 35-37 tok/s on M2 Max via LM Studio.


interestingly, i am also on M2 Max, and i get ~66 tok/s in LM Studio on M2 Max, with the same 131072. I have full offload to GPU. I also turned on flash attention in advanced settings.


Thank you! Flash attention gives me a boost to ~66 tok/s indeed.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: