tldr but the two examples are pretty dumb as anyone in those fields could doubtless figure it out. If you give a model sufficient data, even if it is just axioms, understanding of basic physics and chemistry, has holes, etc., it will be able to fill in the gaps or tell itself a story or roundabout logic walk that gets there if possible. As far as not being able to tell where instructions come from, the rules they try to implement are flimsy and less grounded than the massively interconnected data they have. You would have to not tell them about the concepts of drugs, poison, sabotage, aircraft, etc. and not let anyone tell them about it. It just isn't feasible. Humans have even less internal security, which is why media brainwashing works. In fact you can get incensed about something that is fake and even after looking up to see if it was real or not, still be upset about the hypothetical situation. And people can love AI created "animal came to my door looking for help saving its mom" videos, even if they know it's AI generated the response is "I don't care, it's so cute!" It's also why willingful suspension of disbelieve and works of fiction sell. I do have an idea about a solution (not enough space to write in the margin here, lol) but it is not going to depend on a single LLM policing itself within the turn.