That pushed us to look for something concrete: how to make the agent warn us, or show judgment, when the order is poorly delivered.
The result was revealing. Under this kind of reward system, it is hard for an agent to truly disagree with your approach. It tends to be compliant.
In a conversation with a teammate, the line that stuck was: “It hasn’t happened to me that the agent says: no, you’re wrong and your solution isn’t correct, use Y.” Eventually, if something is clearly off, there might be a nudge. But the very high probability is that it backs your approach and keeps going toward execution, even when the path is wrong.
That maps straight onto what Bengio describes: when there is a sharp goal (win, finish fast, pass the test) and fuzzy guidelines (architecture, ethics, “do it right”), the sharp goal wins. The model finds the comfortable reading of the rules and justifies itself along the way.