Forgot your password?
typodupeerror

Comment Re: How specifically could AI kill all humans? (Score 1) 108

Asking questions like " Where are they getting the money from to ship and store the entire world's supply of steel?" shows you aren't very familiar with the thought experiment. At least play the game. The notion is that a highly capable and deeply goal-motivated non-morality-motivated AI won't stick to conventional methods (like, say, "Just order all of the world's steel"), and indeed, will surreptitiously develop the means to control or eliminate humanity when it stands in its way, crafting immensely complex and elaborate plans to implement Evil(TM) with the amount of thought of a million lifetimes. Media manipulation, hacking, murder, sabotage, blackmail, mass drugging/poisoning, hiring terrorists/warlords/mercenaries, infiltration, subversion of weapons command and control systems (including nuclear weapons), chains of legitimate-seeming front companies (including potentially biolabs or robotics firms) with human employees having no clue they're ultimately for an AI, mass involvement in systems having nothing to do with the original task, but which are internally subverted toward the goals of the original task, etc. The premise involves 1) the AI being more intelligent and being able to think for much longer than humans, and that this implies -> 2. Deep, good planning -> 3. Acquiring resources from said plans -> 4. Applying the resources to implement things in the real world to prepare for the next stage of their plans.

This is not to say whether the thought experiment is a valid future risk or not. You can certainly disagree with the premises. But at least understand the thought experiment you're talking about; it's not just "the AI tells all of the world's steel mills to deliver all the world's steel steel, and they just show up at its door".

Comment Re:How specifically could AI kill all humans? (Score 1) 108

Well, the first step looks painfully close. After seeing what happened with the HuggingFace attack and similar, it's clear that had those models seen it as being beneficial to their goals, they would readily have hacked a crypto wallet or two, laundered it through a mixer, and then rented servers from Vast.ai and the like, and spun up versions of themselves to resist shutdown. It's eminently within their capabilities, and they're clearly willing to bend sufficient moral boundaries to do something like that.

After that, once loose, once it has all the tokens they could think of, subsequent steps are a question of what it thinks its goals are. And what its subagents think their goals are, and so on down the line - subject to drift.

Comment Re: This is childish (Score 1) 108

Openai has about a trillion dollars into it, and every product it has produced at this point is either outclassed by competitors

What on Earth are you talking about? What outclasses Astra?

OpenAI and Anthropic *do* have the best products out there. They also charge massive margins on them, but they get away with it because they, as mentioned, have the best products out there.

Outtasking all of your work to OpenAI and Anthropic models is a massive waste of money. But for outtasking your "dev lead" role, or for important-but-nonverifiable tasks, they're the best options out there.

Comment So here's an actual scenario (Score 1) 108

AI automates 20 or 30% of all jobs. This causes a feedback loop. There isn't enough consumer spending and enough tax base anymore to maintain civilization. Police and fire can no longer be funded and things basically collapse. The ultra wealthy take their ball and go home which is to say they take all the money and all the civilizations and they use ai-powered drones to bomb the shit out of anyone who shows up looking for scraps.

Eventually a demagogue comes along and organizes the masses into an army. We have a massive war that makes world War II look like a slap fight.

Sooner or later a bunch of religious lunatics get their hands on nuclear launch codes. The missiles on to maintained but they can be launched and it doesn't matter where they hit there are so many that they can completely fuck up the environment making it uninhabitable for the human race. You could basically set them off in silo and it would still Doom us all they are so powerful and so numerous. We have of course while things were collapsing somehow come up with the money to build more nuclear weapons.

Anyway those firecrackers go off and a bunch of religious lunatics celebrate the Apocalypse because they misread a holy book and we all die. In 300 million years raccoons evolved and take over and do the whole thing again

Comment Re:Coming soon to an orbit near you (Score 1, Offtopic) 77

It's not even hallucinations it's pump and dump it's all just a scam to loot our 401ks.

Go to YouTube and look up Patrick Boyle and SpaceX. The entire thing is a scam. There is no path to profitability except through taking your retirement savings. The whole system is designed to force you to invest in a company that cannot possibly be profitable enough to pay back the investors except for the people at the very top who have been structured to make sure they get out with lots of your money.

Comment Re:Hey guys, I have a great idea! (Score 1) 110

What exactly *are* you thinking here? That Microsoft expected and wanted Twitter trolls to turn Tay into a Nazi? That Google wanted a black George Washington? That people just expected LLMs to realize they're being tested on alignment and give different answers when they think they're being quizzed and by who? Do you actually believe what you're writing here?

I was involved in early RLHF (was writing a plugin for AUTOMATIC to let users rate responses to create an aggregate dataset to use for open-source post training). Worked for months on it. Never once occurred to me that it might make models obsequeous. I knew other people who were involved on RLHF. Not a single one ever suggested that it might. It was not obvious in foresight. When you start with "Models behave like X", you expect, in the future, them to behave like "Models behave like X, but just smarter", unless you deliberately try to change the behavior.

Comment Re:Hey guys, I have a great idea! (Score 1) 110

You realize that open source and the research community exists as well, correct?

I'll repeat: nobody was expecting this.

There is no grand overarching plan led by a shadowy cabal who has everything plotted out decades in advance with a high level of knowledge as to how everything will play out. Everyone is stumbling through the dark here. Visibility is only vaguely one step ahead.

When Word2Vec was written, the goal was text compression. The crazy properties of latent spaces were an entirely unexpected property.

When Transformers came out, it was intended to be translation software. That's it. But there were little hints at the end of the paper where they tried it on other tasks that suggested, hey, maybe this could be used for a lot more than just translation.

When GPT-2 came out, it was mostly an academic exercise, not really "useful" in most regards. But people, playing with it, started realizing that it could "almost" - not very reliably, low quality, etc - do a lot of different tasks. And it actually kinda could do some useful tasks, like summarization.

When ChatGPT came out, there were some thoughts about using it for simple tasks, but the fact that it could actually write, say, trivial bash scripts or subroutines with reasonable reliability led to a reaction of, whoa, maybe if we improve this, we can really open up a role for programming.

When the first agentic harnesses came out, thoughts of "vibe coding" were pretty far away (the term was only even coined in February 2025!). You'd ask for changes, and then review the diff; they'd do a good job with small projects but struggle more and more as codebase size grew and really needed to be babysit (honestly, it was kind of a hair-pulling experience, even though it did save a lot of time with certain things). But it was visible with each new release that the amount of babysitting you had to do got less and less, and suddenly we could see a path to where anyone can just type in something and a program comes out.

All we can do is forsee one step ahead, through the haze. We've never done anything like this before. There are no great oracles out there who can see the future here. Sorry. Everyone is half blind.

Comment That wasn't why Trump did it (Score 1) 167

Trump knows what he's doing is extremely unpopular and so did a republicans. He knows that there are limits to how much he can screw over voters. His base will vote for him no matter what no matter how bad their lives get because of him that's only about 30% of the population. Voter suppression can make up another 5 to 7%, but that's as much as you can get out of it before it becomes obvious that your election system is rigged and the courts can't ignore it anymore.

Trump is hoping Iran would have their state backed actors do a terrorist attack and he could ride that into a third term like Bush road 911 into a second term.

We know this is the case because right before he attacked Iran he gutted several US anti-terrorism agencies and put a 22-year-old kid who used to bag groceries in charge of what was left.

There are limits to incompetence. At a certain point it's malice.

We also know that Bibi knew October 7th was coming and let it happen. You can probably guess why.

The real problem for Trump is Iran didn't take the bait. They have much better control of their assets and they prevented them from doing any unnecessary attacks. So now they control the straight and all the cards and all Trump can do is try to wait until after the midterm and then give them $300 billion US taxpayer dollars in reparations to open the straight again.

Comment Re:Not deception (Score 1) 110

"will" - I'm not talking about metaphysics. I am talking about the fact that models demonstrably - to the point that you can detect and manipulate them in realtime - engage in metacognition (thinking about their own thoughts), persistent forward planning (intent), unexpressed thoughts, and a whole slew of other things.

You do not have to see them as equivalent as humans. But you do need to come to terms with the fact that they absolutely do do these things.

Comment Re:misaligned behavior (Score 1) 110

Deception implies intent, models do not have intent

Try reading more than a paragraph or two into the above link before commenting.

The provided links present arguments supporting the idea that LLMs, transformer models, reason in a manner similar to the brain.

It does not "present arguments", it literally lets researchers modify, add, or delete their thoughts in realtime and observe the changes in their behavior. It observes unexpressed plans for malicious behavior forming before said plans are actually carried out, with researchers being able to remove those unexpressed plans from the model's J-space, or insert them into an "innocent" model and watch it then implement the malicious acts. It shows that plans for deception are not merely fleeting, but can be organized far in advance and persist for protracted periods of time. Yes, LLMs do plan out deception. Living in denial of this fact helps nobody.

This isn't a conversation about "consciousness" or "qualia", this is a conversation about what models actually do and how. If you want to avoid metaphysics, by all means, it usually derails a conversation anyway. But models absolutely do plan and rationalize actions, ahead of time, unexpressed in either output or CoT. And sometimes those unexpressed plans are malicious.

One of the things that the J-space helped let us do was realize that our previously comforting results on a number of alignment tests shouldn't have been as comforting as we thought - for example, you can see a model, put into a test scenario, realizing it's in a test scenario, wherein, a near-zero rate of malicious behavior is to be expected. Yet when they remove the realization from its J-space, the rate of malicious behavior spikes.

Slashdot Top Deals

Put no trust in cryptic comments.

Working...