Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Not really related but I wonder how people benchmark the effectiveness of skills/agents?

I'm seeing the agent working quite fine with just direct prompting and the agent doing things by itself rather than using skills. Is it better for certain task size?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: