Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

so the coding tasks illicit refusal by anthropic classifiers and instead of updating tasks to do similar things that don’t trigger classifiers their choice is to count those as fails? feels wrong given their stated task list.


I think that's reasonable. The goal of the benchmark is to determine how the model performs on real-world tasks. If real-world tasks trigger Anthropic's classifiers, that's a failure on Anthropic's part.

I feel that choosing tasks that don't trip the classifier would also be a form of bias towards Anthropic.

In case it's not clear, the coding tasks are really benign things; there's nothing security-, health- or biology- related in there. The classifier being tripped is definitely unreasonable.


Indeed; I find it really interesting there is so much disbelief and outright hostility to my post, in these comments. I am an AI Realists rather than AI hype-artist, say it as it is. For me, the evidence is clear that for our own use cases OpenAI, Anthropic models should not be the default choice any longer.

Incidentally, if you want to look at this in more depth - my previous post about the eval framework got like zero response; that's where the results, methodology, code for running the evals, evals are all shared.

Anyone can run this to verify it for themselves

https://github.com/ed-is-ai/featherbench




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: