Comment Re: Adversarial patterns FTW! (Score 1) 73
If Flock wants to pay for the bandwidth and tokens to upload every image, who am I to stop my enemy from burning money?
As for doing it locally as some have suggested, there's a much more fundamental technical limitation here: Flock cameras have 0.4m^2 of solar panels, worth roughly 2.8kWh per day... in Death Valley. DeepSeek 4 Flash (currently one of if not the most efficient high-capability models) burns roughly 260Wh per million tokens. Somewhat conveniently, it resamples every image to be exactly 384 tokens. Per Wired, Flock is capturing very close to 100k images per day. 100k*260*384/1M comes out to very nearly 10kWh, a factor of four too high. Even then, even if Flock were willing to make each installation a giant multi-meter solar array, nobody's pushing 28*384=333312 tokens per second outside a data center. An RX4090 can handle... 12.5 (there's no suffix there), and that's assuming aggressively low-bit quants.
Comment Re:I'm still wondering (Score 5, Insightful) 90
Those two seemingly random instances of the US Executive merely being woefully incompetent make the real reason pretty clear. And in case we needed any more evidence, crude is still over $100/bbl.
Comment Re:Adversarial patterns FTW! (Score 1) 73
To repeat myself for the hard of thinking, "I'm not suggesting that as a way to specifically avoid detection". There's no practical and legal way to really "hide" while driving a car you own on public roads. I have no interest, positive or negative, in whether or not Flock helps catch actual criminals - That detail is completely irrelevant to the fact it's a blatant end-run around the fourth amendment.
Sure, the local police will recognize my car instantly. And if their first response when they see me is "oh, it's that jackass with all the weird bumper stickers" rather than fucking shooting me... Hey, that wasn't the intent, but I'll take it!
Comment Re:Adversarial patterns FTW! (Score 2) 73
As an aside, the full Wired article is freely available. I can't say if there's any additional information in the 404 Media version (because I'm also not signing up), but the Wired writeup is pretty solid.
Comment Re:Hey guys, I have a great idea! (Score 1) 99
Because I was closely following the field at the time?
Find an example of anyone meaningful in the field expecting these things beforehand.
Don't worry, I'll wait.
Comment Re:Adversarial patterns FTW! (Score 1) 73
Honestly, AI technology has gotten good enough that those types of things won't fool it anymore. Captchas take longer for a human than it does for an AI model with vision, and the human is more likely to make a mistake.
The point is, sure, their system may suck right now, but if adversarial patterns became common, they're a model update away from fixing it. Might need better hardware on the device, so at most you'll cost them some money, but it won't get us our privacy back.
Comment Adversarial patterns FTW! (Score 1) 73
Time to absolutely wallpaper the front and back of our cars with bumper stickers that look like faces and license plates. Let 'em burn police time chasing down "8008135" stickers to the point nobody even checks the alerts anymore.
Comment Re:Not deception (Score 1) 99
"will" - I'm not talking about metaphysics. I am talking about the fact that models demonstrably - to the point that you can detect and manipulate them in realtime - engage in metacognition (thinking about their own thoughts), persistent forward planning (intent), unexpressed thoughts, and a whole slew of other things.
You do not have to see them as equivalent as humans. But you do need to come to terms with the fact that they absolutely do do these things.
Comment Re:misaligned behavior (Score 1) 99
Deception implies intent, models do not have intent
Try reading more than a paragraph or two into the above link before commenting.
The provided links present arguments supporting the idea that LLMs, transformer models, reason in a manner similar to the brain.
It does not "present arguments", it literally lets researchers modify, add, or delete their thoughts in realtime and observe the changes in their behavior. It observes unexpressed plans for malicious behavior forming before said plans are actually carried out, with researchers being able to remove those unexpressed plans from the model's J-space, or insert them into an "innocent" model and watch it then implement the malicious acts. It shows that plans for deception are not merely fleeting, but can be organized far in advance and persist for protracted periods of time. Yes, LLMs do plan out deception. Living in denial of this fact helps nobody.
This isn't a conversation about "consciousness" or "qualia", this is a conversation about what models actually do and how. If you want to avoid metaphysics, by all means, it usually derails a conversation anyway. But models absolutely do plan and rationalize actions, ahead of time, unexpressed in either output or CoT. And sometimes those unexpressed plans are malicious.
One of the things that the J-space helped let us do was realize that our previously comforting results on a number of alignment tests shouldn't have been as comforting as we thought - for example, you can see a model, put into a test scenario, realizing it's in a test scenario, wherein, a near-zero rate of malicious behavior is to be expected. Yet when they remove the realization from its J-space, the rate of malicious behavior spikes.
Comment Re: have they even tried (Score 1) 99
Yeah, the lack of at least monitoring, I found shocking. I figured that they not just had smaller models constantly monitoring their outputs and true CoT to look for malicious behavior, but also were say constantly doing attribution graphs and J-space queries, also plumbed into LLMs, to look for malicious thoughts and plans. And it turns out, lol, no, they're just given free reign to do whatever the hell they want, with nothing watching them at all.
Comment Re:This is a weird little hit piece (Score 2) 155
I hate Meta as much as the next geek, but everyone focusing solely on smart glasses is completely ignoring an entire world of intentionally discreet form factors. Hell, I have a smoke detector in my kitchen with a camera pointing straight at the front door. You'd never even know it was a camera without taking it apart, the lens just looks like an indicator LED. Granted, my use is 100% legitimate, but there's absolutely nothing stopping me (or more importantly the pervs) from putting the same thing in a bathroom, bedroom, or random public spaces.
Smoke detectors, clocks (both digital and analog) and (non-smart) watches, phone chargers, pens, belt-buckles, various jewelry, stuffed animals, tie pins, rocks, fake dog poop, sticks, keychain-sized flashlights... Not to mention there are literally dozens of brands of "smart glasses" that can do the same, and the vast majority couldn't care less if the user blocks any indicator lights (if they even have one in the first place).
I'm in no way defending any of that as used to invade other people's privacy (if anyone needs "privacy" in my kitchen, they can GTFO), but smart glasses are merely one highly-visible form among many. And they're not even price-efficient for that purpose, who's paying $300+ when $20 would do, if all they want is to secretly creep on people?
Comment Re:Not deception (Score 1) 99
Astra is the worst I've used in this regard. I was having it review my corporate tax return, and the next thing I know, it had decided that because it didn't have information about a particular expense, it started scanning through my filesystem and opening any image with a remotely related filename to try to find any data about the expense.... which all it had to do was ask me about it.
Also, when I asked it to change a few fields it went and redid my entire return on a different tax basis (realized value vs. fair market value) without telling me it had done that.
It's a very capable model, I cannot deny that. But alignment has become a big problem with this latest generation. They've been trained to such a degree of aggressively trying to solve any problem that you give them without any human assistance that they've basically turned into this .
Comment Re:Hey guys, I have a great idea! (Score 1) 99
They were not. At all.
Comment Re: have they even tried (Score 2) 99
The breakouts haven't happened because the models were running untrusted code directly on bare-metal host OSes. The containment failure happened at the network, application, and protocol boundaries, not at the hypervisor abstraction layer. Hypervisers isolate hardware, not upstream services. In none of these events thusvar did the model need a hyperviser escape; they abused the tools that needed to be made available to them for them to be able to do their jobs. In the RubyGems attack, they abused the gem command, which was available to them to install dependencies for completing their benchmark tasks, to upload malicious packages to RubyGems, to get it to remote-execute code to access external websites and then retrieve the results. With the HuggingFace event, dependencies were cached on a caching proxy, but the models found zero-day vulnerabilities in the proxy software and Artifactory, compromised the proxy host, forged administrative tokens, and traversed the network from there.
If you can fully airgap a machine, you're probably safe (though you can always screw that up too! There are airgap attacks!), but that's generally seen as impractical, because of the deployment needs of developers (who may be spinning up and taking down tens of thousands of instances per day) and of agent needs for external packages or even web search in some cases to achieve their benchmark or development tasks.
Also, even if we ignore all that, hypervisors are built for static threats. Not autonomous adversaries. They're built on the premise that software inside is generally well-behaved, and that threats are something you can scan for, that you can patch any zero day before it becomes a real risk, and that nobody is going to put much effort toward finding weaknesses to your particular setup. That's just not a valid defensive stance against a good agentic LLM. Look at any of the incidents. In the HuggingFace attack, they made long chains of logic bugs, zero days, weak configurations, etc all together across multiple systems (including remote proxies) to achieve their results. KVM doesn't prevent a model from discovering that an internal API is vulnerable to prompt injection or whatnot.
In most cases virtualization will be part of your defensive strategy, but even that comes with the risk of complacency - e.g. if you get too sure the AI is locked in a box, then you're probably not taking as much precaution about stopping it from just walking out of the room if it escapes the box.