Forgot your password?
typodupeerror

Comment Re:You are right (Score 1) 147

the vast, overwhelming usage is to produce non-consenting public recordings of people and leak them to Meta for further data mining.

I wonder where we find stats to resolve our disagreement. (I assert it's a rare case, but I'm pulling that "fact" out of my ass. Alas, I think you're also pulling the "overwhelming usage" claim out of yours, too.)

Comment Re:This is a weird little hit piece (Score 1) 147

does anyone at all really buy into the idea that guys are following around the women of Walmart to video them?

I do buy into it, but I don't buy into any allegations that it's typical or a significant fraction of users.

I know people who own guns but never shot anyone. I know people who have internet access but never pirated any movies. I know people who have very fast cars but don't get speeding tickets. You'd really be surprised at what so many people don't do, even though they could.

Comment Re:Not deception (Score 1) 99

"will" - I'm not talking about metaphysics. I am talking about the fact that models demonstrably - to the point that you can detect and manipulate them in realtime - engage in metacognition (thinking about their own thoughts), persistent forward planning (intent), unexpressed thoughts, and a whole slew of other things.

You do not have to see them as equivalent as humans. But you do need to come to terms with the fact that they absolutely do do these things.

Comment Re:misaligned behavior (Score 1) 99

Deception implies intent, models do not have intent

Try reading more than a paragraph or two into the above link before commenting.

The provided links present arguments supporting the idea that LLMs, transformer models, reason in a manner similar to the brain.

It does not "present arguments", it literally lets researchers modify, add, or delete their thoughts in realtime and observe the changes in their behavior. It observes unexpressed plans for malicious behavior forming before said plans are actually carried out, with researchers being able to remove those unexpressed plans from the model's J-space, or insert them into an "innocent" model and watch it then implement the malicious acts. It shows that plans for deception are not merely fleeting, but can be organized far in advance and persist for protracted periods of time. Yes, LLMs do plan out deception. Living in denial of this fact helps nobody.

This isn't a conversation about "consciousness" or "qualia", this is a conversation about what models actually do and how. If you want to avoid metaphysics, by all means, it usually derails a conversation anyway. But models absolutely do plan and rationalize actions, ahead of time, unexpressed in either output or CoT. And sometimes those unexpressed plans are malicious.

One of the things that the J-space helped let us do was realize that our previously comforting results on a number of alignment tests shouldn't have been as comforting as we thought - for example, you can see a model, put into a test scenario, realizing it's in a test scenario, wherein, a near-zero rate of malicious behavior is to be expected. Yet when they remove the realization from its J-space, the rate of malicious behavior spikes.

Comment Re: have they even tried (Score 1) 99

Yeah, the lack of at least monitoring, I found shocking. I figured that they not just had smaller models constantly monitoring their outputs and true CoT to look for malicious behavior, but also were say constantly doing attribution graphs and J-space queries, also plumbed into LLMs, to look for malicious thoughts and plans. And it turns out, lol, no, they're just given free reign to do whatever the hell they want, with nothing watching them at all.

Comment Re:You are right (Score 1) 147

The two different terms you used kind of look like they apply to different things. "Privacy rapist" appears to refer to the person using the product, but "pervert glasses" appears to refer to the product itself. (Maybe you wanted to call the computer a privacy rapist, but I'm going to ass/u/me that you wanted the more mainstream usage of calling the perv who uses the computer a privacy rapist. Please complain if I got this wrong. I only bring up this pedantry, because I want to distinguish computers from their users.)

Computers can be used for lots of things. They can be used for copyright infringement, for example. But I think Slashdot would be in an uproar if we banned computers due to fact that some people use them to infringe copyrights.

But if someone wants to gargoyle up, such that their cameras are on their face instead of on their lapel or hidden in their clothing or hat, the cameras are objectionable because that user might be recording other people.

If we look at it in terms of the user's intent, instead of worrying about what the equipment is capable of doing, this looks like a pretty inconsistent policy.

That said, I'm not going all the way to saying this view is wrong. As a defender, a potential adversary's capabilities are what's important, and you can't trust anyone's intent. So if you're worried someone might record you, it might really make sense to ban all cameras, not just cameras operated by known pervs.

But why don't we have that attitude about equipment's capability to infringe copyright? Shouldn't we ban all computers, since someone might do the wrong thing with them? I guess it comes down to this: because we aren't the defenders in that situation; MPAA is. Seriously, they really did try to ban betamax due to fears that some people might use it wrong. And they bought a whole law (DMCA) to try to ban lots of uses of computers, which they thought might sometimes be precursors to copyright infringement. So maybe this is a rules-for-thee-but-not-for-me thing.

I'm seeing an attitude toward wearable cameras emerge, that policy should be oriented more toward the the innocent maidens than the drooling gargoyles recording the maidens for later fapping. Well, sure, WHEN I PUT IT LIKE THAT! But what if we describe the gargoyle more fairly? Someone at an in-person office meeting, where maybe they wanna see peoples' names with their faces like they did on Zoom. Or someone recording some cops beating a handcuffed suspect. Someone playing games like Pokemon Go. Someone following a course on a map. Or...

What I'm getting at, is that perv glasses might be useful for more than just perving. Maybe they're really great at perving too, but I can think of plenty of harmless applications for them, and I'm not even very visionary or creative.

I don't like that we're looking down on some tech, rather than its abusive users. I'm not particularly attached to wearable cameras and wouldn't ever likely be in the market for Meta's products (any of Meta products, ever) anyway, so it's easy for me to just shrug and let the fascists have their way on this particular issue.

But I sure as fuck don't want to see this toxic attitude continue to spread to other tech. Because if we allow that, then MPAA is going to ban all computers unless they're running MPAA-certified AI lawyers, always spying on us like the pervs that they are. And even that is just the tip of the iceberg on the smackdowns this attitude will unleash.

Wearable cameras aren't evil.

Now put it on a pole and call it Flock, then it's evi-- no wait, that's a problem with the camera's user too, isn't it?

Comment Re:Not deception (Score 1) 99

Astra is the worst I've used in this regard. I was having it review my corporate tax return, and the next thing I know, it had decided that because it didn't have information about a particular expense, it started scanning through my filesystem and opening any image with a remotely related filename to try to find any data about the expense.... which all it had to do was ask me about it.

Also, when I asked it to change a few fields it went and redid my entire return on a different tax basis (realized value vs. fair market value) without telling me it had done that.

It's a very capable model, I cannot deny that. But alignment has become a big problem with this latest generation. They've been trained to such a degree of aggressively trying to solve any problem that you give them without any human assistance that they've basically turned into this .

Comment Re:Why? (Score 4, Insightful) 37

This is for fun, not "production." If someone hacks up something and got to play with pretty lights for 10 minutes before getting bored and building something else, that's great. And if they learned any engineering, that's a bonus. It's ok to stack your legos however the fuck you want to. I don't fucking care if a kid in 1982 did the "wrong" thing on his TI-99/4A or VIC-20. Been there, done that, got a 40-something year long career out of it despite my youthful unorthodox amateur whims.

Use python on quasi-embedded systems all you want to. Whatever figuratively gets you off your ass, is a good thing.

Comment Re: have they even tried (Score 2) 99

The breakouts haven't happened because the models were running untrusted code directly on bare-metal host OSes. The containment failure happened at the network, application, and protocol boundaries, not at the hypervisor abstraction layer. Hypervisers isolate hardware, not upstream services. In none of these events thusvar did the model need a hyperviser escape; they abused the tools that needed to be made available to them for them to be able to do their jobs. In the RubyGems attack, they abused the gem command, which was available to them to install dependencies for completing their benchmark tasks, to upload malicious packages to RubyGems, to get it to remote-execute code to access external websites and then retrieve the results. With the HuggingFace event, dependencies were cached on a caching proxy, but the models found zero-day vulnerabilities in the proxy software and Artifactory, compromised the proxy host, forged administrative tokens, and traversed the network from there.

If you can fully airgap a machine, you're probably safe (though you can always screw that up too! There are airgap attacks!), but that's generally seen as impractical, because of the deployment needs of developers (who may be spinning up and taking down tens of thousands of instances per day) and of agent needs for external packages or even web search in some cases to achieve their benchmark or development tasks.

Also, even if we ignore all that, hypervisors are built for static threats. Not autonomous adversaries. They're built on the premise that software inside is generally well-behaved, and that threats are something you can scan for, that you can patch any zero day before it becomes a real risk, and that nobody is going to put much effort toward finding weaknesses to your particular setup. That's just not a valid defensive stance against a good agentic LLM. Look at any of the incidents. In the HuggingFace attack, they made long chains of logic bugs, zero days, weak configurations, etc all together across multiple systems (including remote proxies) to achieve their results. KVM doesn't prevent a model from discovering that an internal API is vulnerable to prompt injection or whatnot.

In most cases virtualization will be part of your defensive strategy, but even that comes with the risk of complacency - e.g. if you get too sure the AI is locked in a box, then you're probably not taking as much precaution about stopping it from just walking out of the room if it escapes the box.

Comment Re:Whereas the rest of us believe... (Score 3, Informative) 99

And your reason for why these companies keep experiencing mass resignations, esp. from their safety teams, with people giving up huge amounts of money in order to be able to scream to the press and congress that if they're not stopped they're going to kill us all?

OpenAI

Jan Leike (Former Co-Head of Superalignment) - Resigned in May 2024, posting a viral thread warning that at OpenAI, "safety culture and processes have taken a backseat to shiny products" and that the lab was not prioritizing steering superintelligence.

Daniel Kokotajlo (Former Governance Researcher) - Quit in April 2024, forfeiting ~$1.7M in equity to refuse OpenAI's non-disparagement agreement. Co-organized the A Right to Warn letter, estimating a ~70% chance of catastrophe/extinction from reckless AGI races.

William Saunders (Former Technical Staff / Safety Researcher) - Resigned over safety concerns, signed the Right to Warn letter, and testified before the U.S. Senate in 2024 warning about biological-weapons risks in frontier models and a lack of accountability.

Leopold Aschenbrenner (Former Superalignment Researcher) - Fired in April 2024 after circulating internal security memos; published the 165-page treatise "Situational Awareness," warning of unchecked AGI takeoff, severe national security threats, and espionage vulnerabilities.

Pavel Izmailov (Former Reasoning/Safety Researcher) - Terminated alongside Aschenbrenner; subsequently spoke out publicly regarding insufficient governance, transparency, and safety prioritizations.

Miles Brundage (Former Senior Advisor for AGI Readiness) - Resigned in October 2024, publishing a Substack warning that neither OpenAI nor the world is adequately prepared for AGI.

Gretchen Krueger (Former Policy Researcher) - Resigned alongside Jan Leike in May 2024, posting a public statement calling for institutional accountability, humility, and caution rather than tech hubris.

Jacob Hilton (Former Alignment Researcher) - Left OpenAI over cultural and safety concerns; signed the Right to Warn open letter, warning against a "move fast and break things" approach with frontier AI.

Carroll Wainwright (Former Alignment Researcher) - Resigned over internal safety practices; signed the Right to Warn letter warning about catastrophic risks and suppression of whistleblowers.

Daniel Ziegler (Former RLHF / Alignment Researcher) - Departed OpenAI; signatory of the Right to Warn letter warning of extinction-level and societal risks.

Paul Christiano (Former OpenAI Alignment Team Lead) - Left in 2021 to found the Alignment Research Center (ARC); frequently writes and speaks warning that advanced AI poses a 10%-20%+ chance of human extinction without breakthroughs in control.

Marcus Williams (Agent Monitoring Researcher) - Publicly backed employee warnings, posting that human extinction in the near term is likely (~70% risk) without external regulation or a coordinated slowdown.

Helen Toner (Former OpenAI Board Member) - Voted to oust Sam Altman over safety governance and lack of trust; co-authored an op-ed in Foreign Affairs arguing self-regulation by frontier AI companies is dangerous and unworkable.

Tasha McCauley (Former OpenAI Board Member) - Voted to oust Altman alongside Toner; publicly warned about governance failures and the danger of unchecked corporate control over frontier technologies.

Johannes Heidecke (Former Head of Safety Systems) - Departed during safety reorganizations, expressing concern over structural dissolutions of dedicated safety teams.

Chloé Bakalar (Former AI Ethicist) - Resigned as OpenAI's only dedicated AI ethicist amid company-wide shifts deprioritizing non-commercial ethics research.

Josh Achiam (Former Head of Mission Alignment) - Departed after the company repeatedly reorganized and dissolved its mission alignment teams.

Ilya Sutskever (Co-Founder & Former Chief Scientist) - Spearheaded the board action against Altman over safety concerns; officially left in May 2024 to found Safe Superintelligence Inc. (SSI) to isolate safety research from commercial product pressures.

Google & Google DeepMind

Geoffrey Hinton (Former Google VP & Engineering Fellow / "Godfather of AI") - Resigned in May 2023 specifically to warn the world about existential threats, autonomous systems turning against humanity, and the rapid pace of digital intelligence.

Bilal Chughtai (Former DeepMind AGI Safety Researcher) - Resigned in September 2026, writing on X: "I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome".

Josh Engels (Former DeepMind AGI Safety Team Member) - Resigned in September 2026 to join independent evaluations group METR, warning publicly of "immense harm" within five years.

Ramana Kumar (Former Google DeepMind Researcher) - Resigned and signed the Right to Warn letter, calling out labs for gagging employees with restrictive contracts while pursuing dangerous models.

Neel Nanda (DeepMind Mechanistic Interpretability Researcher / Ex-Anthropic) - Signed the Right to Warn open letter, frequently publishing work highlighting how little developers understand what frontier models are actually doing inside their weights.

Alex Hanna (Former Senior Research Scientist, Google Ethical AI) - Quit in 2022, writing a scathing public resignation letter decrying Google’s toxic suppression of critical ethical research.

Dylan Baker (Former Software Engineer, Google Ethical AI) - Resigned publicly in protest over Google's retaliatory treatment of AI ethics and safety teams.

Blake Lemoine (Former Google Software Engineer) - Fired after publicly voicing ethical alarms about LaMDA’s capabilities and corporate secrecy surrounding model developments.

Meredith Whittaker (Former Google Research Lead) - Organized company walkouts over military AI and ethics; now President of Signal, writing and speaking extensively against Big Tech’s concentrated, unaccountable AI deployment.

Jack Poulson (Former Google Research Scientist) - Resigned over Google’s surveillance and military-adjacent AI projects; now leads Tech Inquiry to track Big Tech defense/AI contracting.

Mo Gawdat (Former Chief Business Officer, Google [X]) - Author of Scary Smart; has given numerous media appearances warning that humanity is creating a dangerous digital deity without adequate control or ethics.

Richard Ngo (Former DeepMind Safety Researcher / Former OpenAI Governance) - Writes extensively on catastrophic misalignment, runaway capability jumps, and the inability of current institutions to govern AGI.

Victoria Krakovna (Google DeepMind Research Scientist) - Co-founder of the Future of Life Institute; regularly publishes research and warnings regarding specification gaming and existential risk from misaligned AI.

Tristan Harris (Former Google Design Ethicist / Center for Humane Technology) - Co-created "The A.I. Dilemma," an influential presentation and essay series warning that runaway commercial generative AI poses an existential threat to democracy and global stability.

Anthropic

Jacob Coxon (Former Pretraining Researcher, Anthropic & OpenAI) - Resigned from Anthropic in September 2026, posting a viral thread decrying both companies for "racing straight to self-improving superintelligence and gambling with our lives".

Joe Benton (Former Safety Research Team Lead) - Left Anthropic in September 2026 to join METR, speaking to the press about escalating dangers as labs prioritize capability over containment.

Evan Hubinger (Intent Alignment Lead) - Backed recent whistleblower statements publicly, posting: "We really do earnestly believe AI could kill all humans!" without coordinated slows or enforceable regulations.

Mrinank Sharma (Former Safeguards Research Team Lead) - Resigned in 2026 with a public letter warning that "the world is in peril," having conducted research on bioterrorism risks and deceptive alignment.

Dario & Daniela Amodei (Anthropic Co-Founders / Former OpenAI VPs) - Originally defected from OpenAI along with ~10 researchers over OpenAI's commercial pivot; Dario has authored essays warning that misaligned AI could take over digital infrastructure in as little as 6 to 12 months. Is currently calling for government regulation and a safety pause on pushing the AI frontier.

Jack Clark (Anthropic Co-Founder / Former OpenAI Policy Director) - Writes the weekly Import AI newsletter, regularly warning about model proliferation, misuse, catastrophic biosecurity risks, and the fragility of current safety benchmarks.

That's just three companies.

There is a widespread feeling within these companies that they're stuck in a race that risks catastrophic consequences for everyone, but can't stop unless everyone agrees to at once, because if they do, then the least scrupulous player will just take over the AI space.

Comment Re:misaligned behavior (Score 5, Informative) 99

Hallucination: model asserts something that's false and that it had no specific reason to believe (for example, gives a URL for something but doesn't bother to check if it's 100% correctly written, or remembers a URL correctly but mixes up its content)

Deception: model tells you something it knows to be false. Yes, they do know when they're deliberately deceiving you (quick 5-minute summary here)

Slashdot Top Deals

Contemptuous lights flashed flashed across the computer's console. -- Hitchhiker's Guide to the Galaxy

Working...