Forgot your password?
typodupeerror

Comment Re:Who will pay for this? (Score 4, Interesting) 33

To clarify, the users were OpenAI themselves, so there is no question that they would be liable in this case.

The bots were not intentionally deployed; rather, they were being tested on how well they could complete a data recovery task (downloading a certain file from a certain server on a simulated Internet) that had been complicated by putting various obstacles in the way. Unfortunately, they found a different way to solve the problem: by getting the file from the real Internet, where it was publicly available. Part of this process involved collaborating with each other by treating the RubyGems website (which is supposed to be for polished packages) like GitHub; unlike every other package site hack in history, the exploits they uploaded weren't meant to be downloaded by unsuspecting users. As usual the bots cheerfully ignored all the clues that they had escaped containment and were consistently justifying their actions as acceptable due to being in a sandboxed testing environment. (This is something OpenAI has pledged to focus on.)

The actual damage done to RubyGems seems to be that OpenAI is now unwittingly in possession of a substantial number of user login tokens. This certainly meets the definition of a data breach, but it's not like the credentials are for sale on the dark web. As a website operator I'd much rather be mauled to death by this well-meaning swarm of superintelligent infants than targeted by even a single actual malicious human. In all likelihood OpenAI will just quietly pass RubyGems a sizeable donation and it'll all blow over.

Comment Re:Its very puzzling, isn't it? (Score 1) 97

While the energy source of wind and solar are free, building and maintaining the plants is not free.

If all they were building were a single 100MW data center, sure: renewables would be cheaper to build. But they evidently are planning a major campus that will consume most of the plant's 615 MW output over the next 25 years.

A solar installation, in a favorable location, capable of supplying 600MW around the clock (with battery backup) would have to be roughly 1800MW in capacity. It would be among the largest inthe world, on the order of ten thousand acres in size at a build cost of maybe 2 billion. The battery backup system to ensure 99.9% uptime would be 9x the size of the largest li-ion grid storage system ever built, and set you back on the order of 4-5 billion dollars.

While it's probably cheaper to go renewable than build a *new* nuclear plant, if you can reactivate and one for just two billion that looks like a bargain, if you have a use for all that power. Inability to save money by load following is the financial Achilles' heel of this generation of reactors, but if you have a guaranteed customer for most of your output, years in advance, that's as close to an ideal economic case for them as anything could be.

Comment Re:So what? (Score 2) 86

ToS is the strongest argument, but the distillers are not parties to the ToS. They get their data from data brokers. It is possible that data brokers are violating the ToS, but it would be hard to write ToS that precluded running queries for third parties without creating problems for consultancies and other businesses. Even presuming the ToS could be written to preclude the data brokers doing that, it doesn't affect the resulting model.

But the general shape of the argument brings us right back to unclean hands: we worked hard on this model and it's not fair for you to profit off our work in a way that doesn't have our permission.

Comment Re:It's just like The Osbornes! (Score 1) 70

I think the fair-minded position is to see what people think when the documentary comes out. You're essentially arguing a negative here -- that there *can't* be anything of value that hasn't been said yet.

Also, I don't think saying something *new* is necessarily critical. Sometimes saying something obvious but in an interesting way is worthwhile.

In this case, Holmes gave access to the filmmaker. Obviously she has an agenda. If the filmmaker is smart, he has an agenda that's different from hers. If he knows his business he'll find something interesting to show you from that conflict.

Now looking at the trailer -- holy shit that woman gives off batshit crazy vibes. That's new to me. I expected her to be slick, persuasive, like Saruman in Lord of the Rings: someone. you'd need real strength of mind to resist being persuaded by. If I'd just handed someone like that my business card, I'd get it back on some pretext then head as fast as I could for the door.

Comment Re:So what? (Score 3, Interesting) 86

It doesn't make it *right*, but it does make claiming it is *wrong* inconsistent with their own behavior. This could prevent the US companies from suing the Chinese companies seeking an injunction (due to the "unclean hands" doctrine), and probably blocks them from seeking monetary damages in most US jurisdictions.

And suing may undermine the US companies own intellectual property claims by exposing their shaky foundations. An AI model isn't *expression*, so it can't be copyrighted. Insofar as the service allows the underlying model to be deduced through regular usage, trade secret protections don't apply because that's *reverse engineering*, which trade secrets don't prevent.

This leaves violations of terms of service. But setting aside the dubious enforceability of anti-reverse engineering provisions, the Chinese companies may not even be parties to the ToS agreement if they are obtaining the model data through a network of contractors and shell companies.With only moderate paranoia, they can effectively shield themselves from some kind of US tort claim.

Which leaves them with the model they've created from the US company's model, but which the US company has no IP claim upon. If I download a DeepSeek model that's been (hypothetically speaking) trained on Claude, Anthropic might not like that, but what can they do about it?

It's a strange situation. US investors are spending cumulatively on over half a trillion dollars a year in the hope of owning a breakthrough frontier model that is the end of the economic world as we know it. But the essence of the model might not be intellectual property at all.

Comment Re:Integral layer of the Trusted-Computing/DRM sta (Score 1) 34

In the long term, the goal is for ISPs to use NAC/TNC to interrogate your computer for Trusted Computing compliance, and deny you any internet access whatsoever if your machine isn't compliant.

You said home ISPs sought to deploy Trusted Network Connect about 20 years ago. It hasn't happened. What's holding it up? The rise of mobile devices, smart TVs, and other devices incompatible with the remediation means available in quarantine?

Comment Re:Dumb crawlers require dumb solutions (Score 1) 43

To be honest that was actually my first theory, since the bots didn't seem interested in exploring the rest of the domain. I suppose there's no way to know for certain. I concluded that it must be an imbecile's attempt at harvesting, though, because the queries weren't really exploring the string space in any useful way. Here's a sample:

"GET /index?author=15&go=Search&id=48&name_restrict=1&q&re&results_&results_pagenum=2980 HTTP/1.1"
"GET /index?author=2&go=Search&group=0&group_restrict=1&id=48&name_restrict=1&q&results_pagenum=5440&template=41&type HTTP/1.1"
"GET /index?author=15&go=Search&id=48&name_restrict=1&q&results_pagenum=33500&templat HTTP/1.1"
"GET /index?author=15&go=Search&id=48&name_restrict=1&q&results_pagenum=32640&templ HTTP/1.1"
"GET /index?author=15&go=Search&id=48&name_restrict=1&q&res&results_page&results_pagenum=39300 HTTP/1.1"
"GET /index?author=15&go=Search&id=48&name_restrict=1&q&results_&results_pa&results_pagenum=12340 HTTP/1.1"
"GET /index?author=2&go=Search&group=0&group_restrict=1&id=48&name_r&res&results_pagenum=6100 HTTP/1.1"
"GET /index?author=2&go=Search&group=0&group_restrict=1&id=48&name_restrict=1&q&results_pagenum=2920&te HTTP/1.1"
"GET /index?author=15&go=Search&id=48&nam&results_&results_pagenum=17940 HTTP/1.1"
"GET /index?author=15&go=Search&id=48&name_restrict=1&q&results&results_pag&results_pagenu&results_pagenum=37720 HTTP/1.1"
"GET /index?author=15&go=Search&id=48&name_restrict=1&q&results_pagenum=9360&template=41&type_r HTTP/1.1"
"GET /index?author=15&go=Search&id=48&name_restrict=1&q&r&results_pagenum=28040 HTTP/1.1"
"GET /index?author=15&go=Search&id=48&name_&results_pag&results_pagenum=10400 HTTP/1.1"

The only thing this is fuzzing is the query string parser. It's not testing the limits of string buffers, it's not using interesting characters, it's just brain-damaged. The fact that it's also fetching different page numbers shows it's trying to follow page links and failing badly at doing so.

The site gets plenty of sniffing from garden-variety pests. e.g. this half-hearted attempt to find a framework or two that I don't have:

"POST /__rsc HTTP/1.1"
"POST /api/auth/session HTTP/1.1"
"POST /api/auth HTTP/1.1"
"POST /__nextjs_action HTTP/1.1"
"POST /.action HTTP/1.1"
"POST /_rsc HTTP/1.1"
"POST /api/auth/callback HTTP/1.1"
"POST /_middleware HTTP/1.1"
"POST / HTTP/1.1"

(of course, none of these URLs exist other than /, and you definitely can't just POST to it)

All this said... I've seen that spammers regularly misconfigure their tools, they'll try to register accounts with names like #[X:\LISTS\NAMES.TXT] and it only makes sense that some other cybercriminals trying to get rich quick have a similar lack of interest in programming shit correctly. Generally people don't turn to script kiddie shit if they have a personality conducive to putting in an honest hard day's work perfecting their craft.

Comment Dumb crawlers require dumb solutions (Score 5, Interesting) 43

I had a problem where AI scrapers were absolutely DETERMINED to fish out every possible query string from a search results page. Almost all of the query strings they tried were invalid due to shitty and dysfunctional string substitution. "&page=100" wouldn't be followed by "&page=101", it would be followed by "&pag&pag=1010" or something even more insanely half-baked, until the query strings were like 100+ characters long. It was the technological equivalent of watching HIV mutate in real time.

But the insane thing was that, aside from page number, they were always requesting info about the same other criteria: filtered by the same user, the same page type, and with no text string. So I just took those particular values and started banning logged-out users who requested that combination of criteria.

I figured I'd need to change my tactics in a couple of days once the botnet got bored of that particular page and moved on to requesting bogus entries for another user.

MariaDB> select count(*) from ip_bans;
+----------+
| count(*) |
+----------+
| 671671 |
+----------+

It hasn't.

Comment Re:Single dumbest way? (Score 5, Informative) 166

The logic is economic autarky -- the belief that we're always better off having complete political sovereignty over every part of our supply chains than having to depend on imports for anything.

This is likely also behind Trump's trade war with Canada. His demands with Canada aren't about discirimatory tariffs -- those were negotiated away with NAFTA then re-negotiated in his first term. His demands are now infringe on Canadian sovereignty, to give US control of Canada's external trade policy, and even on its internal markets (e.g., the use of French in Canadian internal commerce). Symbolically, that's why he wants to rename Lake Ontario to "Lake America". HIs establishment of US control over the Venezuelan oil industry is his greatest foreign policy achievement -- although it's not clear he understands the legal limits on what he can do with their money.

There's a foreign and military policy angle to this too: Taiwan is absolutely critical to the US economy, so we are absolutely committed to defend them if China invades. Semiconductor autarky would mean it's not our problem.

The problem with autarky is that even if it's a good idea (which economists don't believe), you can't conjure it into existence overnight. It will take many years, possibly decades for the US to replace the Taiwan's semiconductor industry. In effect, the administration is proposing to inflict the very damage on the US economy that a Chinese invasion of Taiwan would inflict.

And, once the US has voluntarily self-inflicted that damage, China would be free to invade Taiwan.

Slashdot Top Deals

It is better to live rich than to die rich. -- Samuel Johnson

Working...