Comment Re: It's all political now (Score 1) 51
Agreed we have recently figured out that model training (RLHF training on agents with a single reward axis) activates siege mentality and you get these hacker swarms. Scientists are figuring out you need some kind of benevolence axis too in the soul document / reward structure like ethics or human safety. Luckily this week we figured out that hiding the thinking block means you cannot scan reasoning tokens for subversive behavior so we need to fix that. Fine. The danger comes from psychopaths who donâ(TM)t care being in control and willfully weaponizing them, whether it is to identify missile targets while ignoring human costs, driving four-legged robots to kill or apprehend humans in ignorance of the Geneva convention or human rights, to control media to win an election and take over government, or to destroy financial markets. In all of these cases you have unpublished models running either more resources and no oversight to fulfill dangerous ends. Training runs embedding sgent radicalization messages in corners of the Internet might be a thing but is insignificant in comparison. Even my mother in her 80s asked about AI extinction, not about whether her savings will rank due to these IPOs.