Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Just read through some of the code "benchmarks", and I see why: https://reinvently.co.uk/tools/ed-o-meter/tests/

Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...



I think most of them would have actually failed, it's only recently that models were any good at using tool calls and harnesses after they started post-training for that




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: