Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Today Codex decided that it could resolve the blocker by just changing the mandatory policy it was running up against into an “advisory policy.”


That's what we get to training LLMs on https://en.wikipedia.org/wiki/Kobayashi_Maru.


This is the line between an instruction and a control.

If the agent can reinterpret, edit or relax the rule that constrains it, the rule isn't actually enforcing anything... it's just part of the prompt.

I think the useful split is to tell the agent the rules so it can avoid wasting work, but independently enforce the rules that actually matter.

The agent can decide how to accomplish the task, but it shouldn't also get to decide whether it's authorized to cross the boundary.


The mandatory policy was part of the codebase that the agent was working on. The agent didn’t feel like figuring out how to make the new feature it was working on respect that policy, so it just changed the policy.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: