Forgot your password?
typodupeerror

Comment Seems like variant of a known problem (Score 1) 75

This seems like essentially a variant of the Waluigi effect hypothesized here https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post which had the advantage of a pretty fun name for the situation, and is worth reading. . There's also some related prior work by Ball, Gluch, Goldwasser, Kreuter, Reingold, and Rothblum https://arxiv.org/abs/2507.07341 which suggested that fundamental information issues or computational complexity issues meant that making an LLM AI "safe" from jailbreaks by having a less computationally powerful/less intelligent AI system look over its input would always fail, since one could construct inputs that would be complicated tasks or essentially hidden puzzles that the less powerful AI would fail at but would expand out by the more powerful AI into instructions/jailbreaks. Ball et. al's work is also discussed in detail in this Quanta piece https://www.quantamagazine.org/cryptographers-show-that-ai-protections-will-always-have-holes-20251210/ There's also some similar work by Rao, Choudhury, and Aditya https://arxiv.org/abs/2406.12702 which I haven't read in detail but seems also thematically very similar.

Comment Re:Stupid story (Score 1) 66

It isn't "Gospel" but given the highly detailed timeline, and given that Hugging Face has said explicitly that after OpenAI cooperated with them they are confident that's what happened, it looks like the most likely hypothesis. Alternatives involve OpenAI hacking into multiple other companies in a highly illegal way for extremely unclear gains. And again, Hugging has more details than anyone else, and they are confident that that that happened.

Comment Re:Stupid story (Score 1) 66

Whether you call it thinking or merely prompting is besides the point. The AI was "prompted" if you prefer to accomplish a specific task, performing well on a benchmark. It then responded by escaping a sandbox and hacking Hugging Face. Whether you label it as thinking or merely responding to a prompt, it should still be alarming, in that highly unexpected and genuinely dangerous behavior can occur simply due to being prompted to accomplish a goal. This is exactly the point that people like Yudkowsky and Bostrom were making years ago before any of the LLMs, and people dismissed it as groundless, and said that AIs would not do things like that. Turns out, empirically, they do.

Comment Re:Shit journal - open-access - money grab (Score 1) 84

There are legitimate reasons to criticize this piece. It isn't completely clear whether their 89 data points is a representative sample, and if slightly increasing the inclusion criteria or slightly decreasing them would make this go away. And it isn't completely clear if one should expect something similar to show up in essentially random data simply because two major disasters right next to each other will get treated very likely as a single disaster. There are some other criticisms as well. But being an open-access journal where people pay should not be one of them. This is very common in a variety of STEM fields. I'm lucky enough to be in math where many or most of our open access journals don't have any publication fees, but that's not true in many other fields, and it shouldn't be a reason by itself to discount an article.

Comment Re:"Clean Energy" (Score 3, Interesting) 86

This isn't US specific. South Korea and France have had a lot of successful use of nuclear power. And if you look at different types of power and the negative externalities they produce, nuclear power is very low. See https://janrosenow.substack.com/p/the-bill-we-never-see for a good run down of the different power types and their actual costs.

Comment Re:"dismantle safety protections on Google Play" (Score 1) 69

Hi, I came across this thread and it's something which has raised a lot of concerns for me since Google's announcement. There's a discussion of how it hurts the open-source ecosystem at https://www.reddit.com/r/selfh... and Google's announcement is at https://android-developers.goo...

I like the idea of writing free, open source utilities for others to use. Users having to jump through even more hoops to sideload, or me having to pay money and reveal my identity to Google, get in the way of this --- especially if one of the things I've thought about is e.g. writing a tool to help LGBTQ+ people find locations where they can be reasonably safe in the current US geopolitical climate. That is exactly the sort of utility where the DoJ would make claims I'm some sort of subversive "antifa threat" and demand Google release my identity to them for harassment. Plus, the time where that utility would be most useful -- if someone suddenly has to go someplace unfamiliar on short notice (e.g. getting a call that a relative who lives a few hundred miles away in MAGA territory is getting emergency surgery), having to wait 24 hours to sideload the app if you don't have sideloading currently unlocked, just to find out which rest stops are safe to stop at or not on your drive, would be incredibly unfortunate!

I'm not saying improving app security is a bad thing, and at least Google didn't pull a "We Own Everything And There Is No Way Out Of Our Grip" like they could have, but the options they provide are not nearly good enough. It seems like they only thought about (1) average uninformed users who install things randomly, (2) big commercial developers [who still can get compromised and end up pushing malware, or might even include malware intentionally: just look at LG's McAfee situation, the infamous Sony BMG CD rootkits, etc etc], (3) independent developers who have identity documents Google is willing to recognize [what about the identity document fiascos going on right now with Ukrainian towns Russia is occupying? what about areas occupied in the Middle East by various regimes? what about the way cops confiscate homeless peoples' belongings including ID cards, but you need an ID card to get your belongings back?] plus extra cash [and $25 USD is a lot in some impoverished places, limiting the ability of people in such places to help themselves by becoming app developers], (4) students and hobbyists who only need to share their app with a small group of testers [limiting your ability to test scalability!], and (5) power users who are patient. If you don't fit into the scenarios they planned for, you might be out of luck. And I recognize that the vast majority of legitimate use cases do fall into those above categories, but there are legitimate use cases which don't. They could have at least had a series of open community fora to get feedback on other use cases and plan additional options, but they did not.

It's great that they thought about a few different use cases, at least. But that's not nearly enough!

Could I find an ally to register as the "developer" and publish such an app? Sure, but is that how it *should* be? I would now be dependent on someone else to manage the Play Store listing, which limits my ability to build. Could I spread the word to people who might need such an app that they should unlock sideloading indefinitely just in case? Sure, but maybe that creates new risks for them (and -- what if the option to unlock sideloading indefinitely gets taken away at some future point? Google gets to dictate all the changes here, and users have little recourse if they disagree with the specifics of the implementation)

I just want the guaranteed ability to say "hey, what about XYZ situation, which I or a friend have personal experience with? can you provide a way to handle that?" but I don't. There's no form to fill out to request some sort of special accommodation for an unusual situation, and there probably should be. When you have billions of users, planning for the needs of 99.99% of them still hangs hundreds of thousands of users out to dry. I know trying to accommodate more edge cases isn't as profitable, and having a large enough support team to hear people out when they have an unusual need requires a ton of extra staffing, but smartphone usage has become a critical utility rather than a luxury in much of the world. Google has an immense amount of power over numerous aspects of the lives of all sorts of different people, and if they want to continue to make money from that situation, they need to be prepared to deal with the complexities of it all.

Slashdot Top Deals

"Time is money and money can't buy you love and I love your outfit" - T.H.U.N.D.E.R. #1

Working...