Packet filtering, red teams, and all other security cyberdung is simply not interesting here. Just assume, for a moment, that OpenAI guys not only not worse IT pros than you but actually very much better. I understand that it is hard to believe, but do make that effort just for the sake of your own understanding.
The problem belongs to a higher level of abstraction. Assume that OpenAI built a perfect network isolation. Nothing, absolutely nothing gets outside of the lab unless OpenAI decide to let it through. To train those hacking AI algorithms, OpenAI still must grant them access to the Internet (that access must be granted promptly, AI-training promptly, otherwise the algorithms either cannot train at all or they do it so slowly that the competition takes over). Granting access, on the other hand, is dangerous.
This is the issue at hand. How do you decide whether to push that knob granting access? Forget all that quasi-security jargon. Imagine that you have a single button in front of you that is labelled LET IT THROUGH. The button functions perfectly. No push — not a single octet gets out. (Note that I am helping you by modelling a proper whitelisting mechanism.) When do you push it and when you ignore a request to push it?
To deepen your understanding of the issue, imagine that the algorithm may figure out that not all of its requests make it through. Since this is not a human conscience, the algorithm turns into a bitter liar attempting to masquerade its suspicious requests, turning more skilful every day.
Continue thinking in that direction and, perhaps, you shall be enlightened.
It is not an IT problem or a security problem.