Comment What they think is not what we really need. (Score 1) 97
AI neural models aren't programmed in a classic sense, but trained.
That means that any rule works OVER the true neural model that make the reasoning or decisions.
Yes... You can also control the data of the training. But that doesn't guarantee a way of reasoning or future deviations with evolving models that catch input constantly.
So the rules are A JAIL, not an intrinsic way of reasoning of the model. In other words, the own model can search ways to overcome its own jail.
What you want is an AI that recognized that act as bad, NOT one that decides certain course of action, and an external verification software layer crops that result become it detects is a damaging/banned action.
On other words... we need to UNDERSTAND how the model build the answers to ensure that answer is recognized as bad just at the start. In summary, to ensure the AI have a kind of "strict moral code" already trained in the deep level of the model.