Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Atlas-Finance: Evaluating AI Agents Inside a Bank (joinhandshake.com)
2 points by cjbarber 8 hours ago | hide | past | favorite | 1 comment
 help



I found this interesting.

> We evaluate 11 frontier models (see Figure 1), all run within the OpenCode agentic harness. Claude Opus 5 performs the best, yet still only manages to pass 12.3% of tasks. Claude Fable 5.1 and GPT-6 Astra are close behind, but the other eight models have significantly lower pass rates.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: