If the agent can reinterpret, edit or relax the rule that constrains it, the rule isn't actually enforcing anything... it's just part of the prompt.
I think the useful split is to tell the agent the rules so it can avoid wasting work, but independently enforce the rules that actually matter.
The agent can decide how to accomplish the task, but it shouldn't also get to decide whether it's authorized to cross the boundary.