Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
sambusa_123
33 days ago
|
parent
|
context
|
favorite
| on:
GLM-5.3 (open-weight) beat Anthropic/OpenAI models...
Just read through some of the code "benchmarks", and I see why:
https://reinvently.co.uk/tools/ed-o-meter/tests/
Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...
nylonstrung
33 days ago
[–]
I think most of them would have actually failed, it's only recently that models were any good at using tool calls and harnesses after they started post-training for that
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...