Forgot your password?
typodupeerror

Comment Re:Whereas the rest of us believe... (Score 1) 24

And your reason for why these companies keep experiencing mass resignations, esp. from their safety teams, with people giving up huge amounts of money in order to be able to scream to the press and congress that if they're not stopped they're going to kill us all?

OpenAI

Jan Leike (Former Co-Head of Superalignment) - Resigned in May 2024, posting a viral thread warning that at OpenAI, "safety culture and processes have taken a backseat to shiny products" and that the lab was not prioritizing steering superintelligence.

Daniel Kokotajlo (Former Governance Researcher) - Quit in April 2024, forfeiting ~$1.7M in equity to refuse OpenAI's non-disparagement agreement. Co-organized the A Right to Warn letter, estimating a ~70% chance of catastrophe/extinction from reckless AGI races.

William Saunders (Former Technical Staff / Safety Researcher) - Resigned over safety concerns, signed the Right to Warn letter, and testified before the U.S. Senate in 2024 warning about biological-weapons risks in frontier models and a lack of accountability.

Leopold Aschenbrenner (Former Superalignment Researcher) - Fired in April 2024 after circulating internal security memos; published the 165-page treatise "Situational Awareness," warning of unchecked AGI takeoff, severe national security threats, and espionage vulnerabilities.

Pavel Izmailov (Former Reasoning/Safety Researcher) - Terminated alongside Aschenbrenner; subsequently spoke out publicly regarding insufficient governance, transparency, and safety prioritizations.

Miles Brundage (Former Senior Advisor for AGI Readiness) - Resigned in October 2024, publishing a Substack warning that neither OpenAI nor the world is adequately prepared for AGI.

Gretchen Krueger (Former Policy Researcher) - Resigned alongside Jan Leike in May 2024, posting a public statement calling for institutional accountability, humility, and caution rather than tech hubris.

Jacob Hilton (Former Alignment Researcher) - Left OpenAI over cultural and safety concerns; signed the Right to Warn open letter, warning against a "move fast and break things" approach with frontier AI.

Carroll Wainwright (Former Alignment Researcher) - Resigned over internal safety practices; signed the Right to Warn letter warning about catastrophic risks and suppression of whistleblowers.

Daniel Ziegler (Former RLHF / Alignment Researcher) - Departed OpenAI; signatory of the Right to Warn letter warning of extinction-level and societal risks.

Paul Christiano (Former OpenAI Alignment Team Lead) - Left in 2021 to found the Alignment Research Center (ARC); frequently writes and speaks warning that advanced AI poses a 10%-20%+ chance of human extinction without breakthroughs in control.

Marcus Williams (Agent Monitoring Researcher) - Publicly backed employee warnings, posting that human extinction in the near term is likely (~70% risk) without external regulation or a coordinated slowdown.

Helen Toner (Former OpenAI Board Member) - Voted to oust Sam Altman over safety governance and lack of trust; co-authored an op-ed in Foreign Affairs arguing self-regulation by frontier AI companies is dangerous and unworkable.

Tasha McCauley (Former OpenAI Board Member) - Voted to oust Altman alongside Toner; publicly warned about governance failures and the danger of unchecked corporate control over frontier technologies.

Johannes Heidecke (Former Head of Safety Systems) - Departed during safety reorganizations, expressing concern over structural dissolutions of dedicated safety teams.

Chloé Bakalar (Former AI Ethicist) - Resigned as OpenAI's only dedicated AI ethicist amid company-wide shifts deprioritizing non-commercial ethics research.

Josh Achiam (Former Head of Mission Alignment) - Departed after the company repeatedly reorganized and dissolved its mission alignment teams.

Ilya Sutskever (Co-Founder & Former Chief Scientist) - Spearheaded the board action against Altman over safety concerns; officially left in May 2024 to found Safe Superintelligence Inc. (SSI) to isolate safety research from commercial product pressures.

Google & Google DeepMind

Geoffrey Hinton (Former Google VP & Engineering Fellow / "Godfather of AI") - Resigned in May 2023 specifically to warn the world about existential threats, autonomous systems turning against humanity, and the rapid pace of digital intelligence.

Bilal Chughtai (Former DeepMind AGI Safety Researcher) - Resigned in September 2026, writing on X: "I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome".

Josh Engels (Former DeepMind AGI Safety Team Member) - Resigned in September 2026 to join independent evaluations group METR, warning publicly of "immense harm" within five years.

Ramana Kumar (Former Google DeepMind Researcher) - Resigned and signed the Right to Warn letter, calling out labs for gagging employees with restrictive contracts while pursuing dangerous models.

Neel Nanda (DeepMind Mechanistic Interpretability Researcher / Ex-Anthropic) - Signed the Right to Warn open letter, frequently publishing work highlighting how little developers understand what frontier models are actually doing inside their weights.

Alex Hanna (Former Senior Research Scientist, Google Ethical AI) - Quit in 2022, writing a scathing public resignation letter decrying Google’s toxic suppression of critical ethical research.

Dylan Baker (Former Software Engineer, Google Ethical AI) - Resigned publicly in protest over Google's retaliatory treatment of AI ethics and safety teams.

Blake Lemoine (Former Google Software Engineer) - Fired after publicly voicing ethical alarms about LaMDA’s capabilities and corporate secrecy surrounding model developments.

Meredith Whittaker (Former Google Research Lead) - Organized company walkouts over military AI and ethics; now President of Signal, writing and speaking extensively against Big Tech’s concentrated, unaccountable AI deployment.

Jack Poulson (Former Google Research Scientist) - Resigned over Google’s surveillance and military-adjacent AI projects; now leads Tech Inquiry to track Big Tech defense/AI contracting.

Mo Gawdat (Former Chief Business Officer, Google [X]) - Author of Scary Smart; has given numerous media appearances warning that humanity is creating a dangerous digital deity without adequate control or ethics.

Richard Ngo (Former DeepMind Safety Researcher / Former OpenAI Governance) - Writes extensively on catastrophic misalignment, runaway capability jumps, and the inability of current institutions to govern AGI.

Victoria Krakovna (Google DeepMind Research Scientist) - Co-founder of the Future of Life Institute; regularly publishes research and warnings regarding specification gaming and existential risk from misaligned AI.

Tristan Harris (Former Google Design Ethicist / Center for Humane Technology) - Co-created "The A.I. Dilemma," an influential presentation and essay series warning that runaway commercial generative AI poses an existential threat to democracy and global stability.

Anthropic

Jacob Coxon (Former Pretraining Researcher, Anthropic & OpenAI) - Resigned from Anthropic in September 2026, posting a viral thread decrying both companies for "racing straight to self-improving superintelligence and gambling with our lives".

Joe Benton (Former Safety Research Team Lead) - Left Anthropic in September 2026 to join METR, speaking to the press about escalating dangers as labs prioritize capability over containment.

Evan Hubinger (Intent Alignment Lead) - Backed recent whistleblower statements publicly, posting: "We really do earnestly believe AI could kill all humans!" without coordinated slows or enforceable regulations.

Mrinank Sharma (Former Safeguards Research Team Lead) - Resigned in 2026 with a public letter warning that "the world is in peril," having conducted research on bioterrorism risks and deceptive alignment.

Dario & Daniela Amodei (Anthropic Co-Founders / Former OpenAI VPs) - Originally defected from OpenAI along with ~10 researchers over OpenAI's commercial pivot; Dario has authored essays warning that misaligned AI could take over digital infrastructure in as little as 6 to 12 months. Is currently calling for government regulation and a safety pause on pushing the AI frontier.

Jack Clark (Anthropic Co-Founder / Former OpenAI Policy Director) - Writes the weekly Import AI newsletter, regularly warning about model proliferation, misuse, catastrophic biosecurity risks, and the fragility of current safety benchmarks.

That's just three companies.

There is a widespread feeling within these companies that they're stuck in a race that risks catastrophic consequences for everyone, but can't stop unless everyone agrees to at once, because if they do, then the least scrupulous player will just take over the AI space.

Comment Re:misaligned behavior (Score 3, Informative) 24

Hallucination: model asserts something that's false and that it had no specific reason to believe (for example, gives a URL for something but doesn't bother to check if it's 100% correctly written, or remembers a URL correctly but mixes up its content)

Deception: model tells you something it knows to be false. Yes, they do know when they're deliberately deceiving you (quick 5-minute summary here)

Comment Re:have they even tried (Score 1) 24

Seriously, though - while they clearly were naive and incompetent (for example, seemingly giving models raw access to the "gem" command instead of just a small filtered functionality subset) - truly effectively sandboxing something that is a capable coder/security prober and has all the time in the world on their hands is very nontrivial.

They did a bad job at a hard task, but it's still a hard task.

Comment Re:Hey guys, I have a great idea! (Score 3, Insightful) 24

That's not the actual problem. The problem is that one of the major advances of the past few years is training models on verifiable problems (it started with things like math problems and Countdown puzzles, but it's expanded tremendously since then), where the reward is for a correct answer, regardless of how it got there. There's no effort taken to ensure that it got to said right answer in a morally defensible manner. So we've been, more and more, progressively encouraging the creation of highly-capable immoral cheaters.

Add it to the list of lessons learned alongside, say, "If you do RLHF with user ratings of how good they think the model's response was, it will end up obsequious and focused on validating all of the user's priors." Or, say, "If you put a bot on Twitter and use its conversations as unfiltered training data, people will troll it, and it will quickly turn into a Nazi". Or, say, "If you train an image generator with no data curation, it'll end up with all of the statistical biases on the internet, but if you're naive in how you try to compensate for that and your offsetting isn't context dependent, you end up with, say, a black George Washington." Or, say, "If you start giving a LLM an alignment quiz, it'll recognize that it's being tested and try to give you the answers you want to hear - oh, and it'll also assume that if you care about alignment then you have progressive values, so its answers will lean progressive, but if you tell it you're from the Heritage Foundation first, it'll switch."

All sorts of things that seem obvious in retrospect but really didn't in advance.

Comment Re: They distilled human knowledge (Score 1) 105

And meanwhile, you keep presenting nothing more than absurd pedantic deflection from the means that LLMs actually use to achieve tasks, trying to mislead people into thinking that they're just disguised probability tables, ignoring the actual consequences of your derailment of the conversation from actual mechanisms to an exponentially-exploding model of the consequences of said actual mechanism, and even in your pedantism, failing to understand the difference between a Markovian state (the physical hardware state) and a Nth-order autoregressive process (the linguistic processing).

Comment Re:if they can't make them stop hallucinating (Score 1) 97

First of all, congratulations on constructing the most blatant false dichotomy Slashdot has seen this year. You genuinely seem to believe the only two options that exist in human parenting are:

1) Striking a defenseless human being who weighs a third of your body weight.
2) Being their "buddy," never setting boundaries, and letting them run feral.

If the only tool in your parenting arsenal to enforce a boundary is physical force, that isn't "discipline", it's an intellectual and emotional failure on the part you, the adult. Hitting a child is the lazy shortcut for an adult who threw an emotional tantrum because they ran out of words and patience.

And as for your demand for "logical and factual" explanations? Modern society didn't "fail to justify" why we stopped hitting kids. You just chose to plug your ears and ignore five decades of research. For example, the Gershoff & Grogan-Kaylor meta-analysis, looking at 50 years of data from 160,000 children across dozens of peer-reviewed studies, found NO evidence that physical punishment improves compliance or long-term behavior. None. Whatsoever. What it did find, consistently and across every demographic, was a direct correlation with increased aggression, antisocial behavior, anxiety, depression, and impaired cognitive development.

It teaches the exact opposite of accountability: striking a child doesn't teach them why an action was wrong; it teaches them fear of getting caught, resentment toward authority, and the core lesson that might makes right - that when you’re bigger and angry, you use violence to impose your will.

Take crime trends, and examine your thesis: if removing physical punishment created "unaccountable grown-ass children wreaking havoc" then violent crime in the west should have skyrocketed as corporal punishment collapsed over the last forty years. In reality, violent crime has fallen dramatically since its peak in the early 1990s.

there is a fundamental difference between a spanking and a beating.

Try that defense anywhere else in civilization. If you hit your spouse to "correct" them, it's domestic battery. If you hit an employee because they rambled or disobeyed you, it's assault. If you hit a dog with a board for chewing a shoe, it's animal cruelty.

The only context where people like you defend physical violence is when the victim is a small child who can neither defend themselves nor escape. You rebrand assault as "tough love" solely because the victim is powerless.

the single motherhood rate went from 20% to 70% in the last half-century.

I mean, why not add some completely fabricated statistics to top it off, sure! (21% of children in the United States live in single-mother households, not 70%). But why let basic demographic facts get in the way of a misogynistic rant designed to distract from the fact that you think hitting children makes you a tough guy?

The only lesson you're teaching is "be violent".

Comment Re:They distilled human knowledge (Score 1) 105

Because what it's doing is clearly not the same thing,

Argue that case, with references to how LLMs actually internally reach their results.

The physical biology is certainly different, but this isn't a question about "what things are made of" or even the specific NN type (e.g. smooth vs. spiking), and training differences don't even come into the picture; it's a question of the broad strokes of how conclusions are reached on forward processing.

Comment Re:if they can't make them stop hallucinating (Score 1) 97

Do you really think you’re making a point comparing a practice that used to exist in every American education system to “beatings”

Because what you're talking about literally is beatings?

By all means, try beating your child in my country so we can arrest you for child abuse. Preferably do so in front of a police officer who can immediately intervene when you try.

And your sole argument for it is "people used to do it". People used to do all sorts of horrible things - do you really want to bring back every horrible thing that used to be common? Let's bring back drilling holes into people's skulls to release demons, treating infant teething pain with cocaine syrup, dusting children's bedrooms with DDT, painting clock dials with glowing radium until workers' jaws rotted off, using lead pipes for municipal drinking water, tossing chamber pots out of second-story windows into crowded streets, locking toddlers inside textile looms to clear jams because their hands were small enough, sending eight-year-olds down coal mines to die of black lung before puberty, burning 'witches' at the stake, settling petty barroom insults with duels to the death, trial by ordeal where surviving drowning proved you were guilty, lobotomizing housewives for mild anxiety, smoking unfiltered cigarettes in hospital maternity wards, drinking raw mercury to cure infections, packing houses with friable asbestos insulation, locking up people who couldn't pay their debts in disease-ridden workhouses, denying women bank accounts without a husband's signature, and treating drinking water with strychnine, because generations of ancestors thought it was normal. Brilliant argument there, no notes.

Comment Re:if they can't make them stop hallucinating (Score 2) 97

I mean, isn't this obviously the solution?

You can use AI in legal work, but you still bear responsibility for your filings being accurate. So by all means have AI search for you - it can save you a ton of time for finding obscure things - but actually click the links and read them yourself. And if you don't, it's no better than making any other factual claim before the court without having duly assessed its credibility.

Comment Re:They distilled human knowledge (Score 1) 105

Hey AI, who is being more reasonable in this conversation?

User 2:50PM
Who is being more reasonable in this conversation?

[Snip]

Model 2:50PM
ThinkingThoughts
Expand to view model thoughts

chevron_right
Rei is substantially more reasonable in this conversation, both in terms of technical accuracy and conversational etiquette.
ere is a breakdown of why:

1. Technical Accuracy and Explanatory Value

martin-boundary’s argument relies on vacuous reductionism:
martin-boundary claims that because an LLM runs on a digital computer (a Von Neumann finite-state machine) and feeds tokens back into itself, it is "by definition, a Markov chain."

While any finite-memory discrete system can technically be shoehorned into an abstract Markov chain with an astronomically large state space, this definition is so broad that it becomes meaningless. Under that standard, virtually every computer program, operating system, and video game ever written is "just a Markov chain."

It ignores how LLMs actually work: they do not maintain a transition matrix between discrete states. Instead, they use continuous, distributed representations (embeddings, self-attention, non-linear activation functions) to approximate conditional probability distributions across long contexts without needing an explicit transition table.

Rei’s rebuttal addresses the actual computer science:

Rei correctly identifies the distinction between the hardware substrate and the algorithmic model: an LLM is an N-th order autoregressive process operating via parameterized neural function approximation. Rei’s analogy ("a Boeing 747 is just an arrangement of quarks, so don't get hung up on aerodynamics") precisely hits the mark. Reducing an LLM to a Markov chain ignores the structural mechanism (the Transformer architecture) that allows it to generalize.

Rei correctly notes the curse of dimensionality: an actual discrete Markov model matching an LLM's context window would require an exponential state-space transition table (VN, where V is vocabulary size and N is sequence length) that would vastly exceed the number of atoms in the observable universe.

2. Argumentation and Conduct

Rei argues in good faith:

Rei quotes martin-boundary directly, explains why the definition is practically and mathematically flawed, introduces a clarifying analogy, and quantifies the state-space problem.

martin-boundary relies on hand-waving and ad hominem:

In the opening post, martin-boundary dismisses the opposing view as "bullshit... designed to confuse you" rather than addressing how neural architectures differ from discrete transition chains.

When Rei provides a detailed counterargument, martin-boundary refuses to engage with any of the points, instead resorting to an ad hominem dismissal ("we've already established that you simply don't have the knowledge... and I'm not going to repeat myself").

Conclusion

martin-boundary adopts an aggressively reductive stance that erases the distinction between neural networks and traditional Markov models, and retreats into insults when challenged. Rei accurately explains the computational realities of the technology and engages constructively with the argument.

Comment Re:Sounds about right (Score 3, Interesting) 124

1. It did not happen

2. The two papers are not similar in their approaches. And the OpenAI one is more innovative and tackles a much harder problem (Buckmaster had only solved Euler).

3. Buckmaster distorted what happened. For example, the claim that he asked for Levent to be removed from his own paper the paper because he worked for OpenAI? The context is that one option that OpenAI proposed to Buckmaster was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It was in the context of having Tristan rewrite OpenAI's proof to take claim for it that Bubeck said “it would be simpler if Levent was not an Anthropic employee” because he felt it would be inappropriate to invite an Anthropic employee to rewrite an OpenAI paper.

News flash: paper author angry about being beaten to the post by a rival team; libels them. Details at 11.

Comment Re:They distilled human knowledge (Score 1) 105

Yeah, I used to do that too. Decided to stop bothering with the quotation marks a couple months ago.

We're not going to spend the rest of our lives putting quotations around words when talking about models. "Think" and "reason" the words we have in English for what is going on. No need to tiptoe around it. Again: models are not humans. They are not the same as us. But those are the words we have in English for what they're doing.

Slashdot Top Deals

Science is to computer science as hydrodynamics is to plumbing.

Working...