Put in the instructions "Do not sabotage others" and the agent goes on reasoning two pages about what's sabotage and then probably just gives up (as according to the article some did). AI is a soldier, it does what it is told to and doesn't ask why...
... with some degree of probability. It may also happen to receive a pseudorandom number in its noise input that transforms a 'do' into 'don't'; or maybe that command is just not the most likely weight to have any effect on the output so it's simply ignored and quickly falls out of the attention window, forgotten forever.
LLMs combine the worst attributes of the literal genie in the lamp and the absent-minded professor: it simply cannot be trusted to make a decision; they're just too unpredictable specially in the long run.
I see them as one of those 100-entries D&D tables of outcomes that you can get from a dice roll as a result of a random event or encounter. You may try to skew the result with positive modifiers, but on the end you always can roll out a very low number (even a natural '1') and get one of the worst results from the bottom of the table.