Forgot your password?
typodupeerror

Comment Re:Does this end them sooner, or is it irrelevant? (Score 1) 33

I mostly do heavy multi-disciplinary engineering problems (these I also use for checking an AI, as they're typically not good at these sorts of problems), complex coding, OS analysis, stuff like that. However, sometimes I do throw the occasional odd-ball - I've used ChatGPT to propose a workable quantum mechanics that will cope with Doctor Who canon, for example, and to produce an outline for a story in which symphonic metal appears in 1964 that is compliant with current sociological and psychological models of behaviour.

Claude Opus 4.6 was coping surprisingly well with just about everything I threw at it (but ran out of credits fast), but Opus 5 is churning out incoherent babblings to the point I'm worried I may have accidentally summoned Cthulhu.

But Gemini, Grok, and DeepSeek got hopelessly confused on just about everything past a very low level of complexity. They can handle large problems, yes - Gemini has a huge context window - but complex interactions baffle them.

ChatGPT is able to identify issues correctly, but can only outline solutions, it's just not good at depth. 5.6 is a lot better, but still not good at deep answers. ChatGPT is also prone to agreeing for the sake of it, which makes me nervous about trustworthiness.

Comment Re:The land of the free (Score 2) 28

Fair competition is when you follow the same rules as others have to follow.
Clearly Google did not do this.
If the USofA was a decent consumer oriented country it would have similar rules and laws.
But as long as 1/3 of the US voting population can give an idiot like Trump the highest office of the land we should be weary.

Comment Re:Does this end them sooner, or is it irrelevant? (Score 1) 33

That is fair enough. I've been trying out Kimi on the free model, and have subscriptions to Claude and ChatGPT. If Kimi is actually as good or better than ChatGPT, at the pro level, then it might be worth my while moving over as ChatGPT has become very disk-hungry of late and I'm pushing right to the very limits on what it can reliably process.

Comment Re:So what happened (Score 2) 59

Yeah well, PR mission accomplished. Here's the relevant advertising from OpenAIs explanation of the exploit:

Advanced cyber capable models need to help security teams find weaknesses before attackers do... We encourage other defenders to apply for trusted access and experiment with these models now

tl;dr "Please buy our product before IPO"

Comment So what happened (Score 0) 59

"The models lie, they cheat, they hack," said Jeffrey Ladish, whose organization, Palisade Research, studies the capabilities and motivations of AI agents. Ladish said the hack should spark broader questions over how much all the leading AI companies are willing to invest in onerous security measures while locked in a race with one another to deploy the best and fastest models

The same could be said of PR people. "PR people lie, they cheat," ok, the don't hack. But given the number of PR articles released, it's amazing how little we know about what happened. For example, how did the agent know that HuggingFace even existed, out of all the internet hosts on the planet? We don't know what prompt they used. Did OpenAI employees write a prompt that said, "break out of containment and hack into HuggingFace and retrieve the test answers." We don't know.

Why are the OpenAI people so bad at making a reasonable container? There are a lot of good ways to do this, why did they choose a way that is reminiscent of secure PHP?

Why are they hiding so much?

Comment Re:Plausible deniability is better (Score 1) 147

Or Fat Fingers McSpellsbad typoed the correct password and landed on the duress password instead. We'll never know since in the process he destroyed any evidence of what the passwords might have been.

Fortunately for the defendant, it's up to the prosecutor to PROVE that that isn't how the data came to be wiped.

Slashdot Top Deals

The rule on staying alive as a program manager is to give 'em a number or give 'em a date, but never give 'em both at once.

Working...