Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
hrpnk
on Aug 5, 2025
|
parent
|
context
|
favorite
| on:
Open models by OpenAI
gpt-oss-20b: 9 threads, 131072 context window, 4 experts - 35-37 tok/s on M2 Max via LM Studio.
rt1rz
on Aug 5, 2025
[–]
interestingly, i am also on M2 Max, and i get ~66 tok/s in LM Studio on M2 Max, with the same 131072. I have full offload to GPU. I also turned on flash attention in advanced settings.
hrpnk
on Aug 6, 2025
|
parent
[–]
Thank you! Flash attention gives me a boost to ~66 tok/s indeed.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: