Forgot your password?
typodupeerror

Comment replication (Score 1) 176

Small sample, but still, food for thought.

No, not at all. The gender pay gap is in the training data thousands of times, from online comments, studies, news articles, etc. Since AI is a dumb machine replicating what it has in its training data, it replicates that.

I'm pretty sure you can find religious believe, superstitions of all kinds, human fallacies and whatever in AI if you provide it with the right setup.

Comment Translator (Score 2) 41

I ran this through the translator. "Having a dedicated home with Prime Video allows us to build on that legacy with unprecedented global reach and a broader commitment to the Television Academy and our mission throughout the year."

It came out "Jeff offered more money."

Comment Re:How many more (Score 2) 106

No one saw quasi-RH coming or many of the others here. So that is unlikely. And the last batch of 10 problems they released a few weeks ago which were impressive enough were all novel results. There have been issues with them using techniques or ideas that should be better credited to where some of those approaches are coming from. But that's the sort of thing that often happens at a preprint stage. When I referee a paper, if this happens, you just request they cite the relevant papers and move on. This isn't at all stealing things.

Comment Re:mostly mash-ups of existing techniques (Score 5, Interesting) 106

If the mathematicians ever used ChatGPT to talk about math, then that is the source of the results. LLMs, by definition, cannot usefully contribute to anything that requires more than rearranging existing data. Rearranging existing data is their only function.

This is really not accurate. There's Fields Medal level work here with quasi-RH for example. And for Hadwiger-Nelson it took an approach that doesn't seem to be in in the literature. These things really are doing novel math. I've personally seen this is in a bunch of situations. Here's a personal example, much smaller than anything like these problems. I have a recent preprint with a student here https://arxiv.org/abs/2609.36068 (about 90% of this was done by her. She's very good.) But part of this came from when I ran a version of Proposition 13 in that paper through Claude just to clean up the draft of that bit before I sent it to her. Claude informed me (essentially unprompted) that the argument had a whole in it (in addition to pointing out grammar errors, unbalanced parentheses and some other embarrassing minor mistakes). I then fixed the hole, and gave it back to Claude. Claude thought for a few minutes, and then informed me that I had *not* fixed the hole, and the reason was that my version of the proposition was missing an entire infinite family which it constructed. The literature on this problem is small, and I'm very familiar with it. The family it produced is straightforward (see Prop 8), but definitely was not in the existing literature. So yes, these systems really can do novel math, and can even do so in an essentially minimally prompted fashion, in this case, explaining to the meat mathematician why his proof is wrong. My student and I then generalized Claude's family to Theorem 9 in that paper, but the AI definitely had an impact.

At some level this was actually not a good thing. I'm trying to encourage students to *not* rely on the AI for research so they develop basic research skills. So if the AI had not volunteered the family I wouldn't have had to tell her that the AI had discovered it, but honestly required noting it. So I had a conflict between intellectual honesty and being a good role model.

Comment A few comments on the problems (Score 5, Informative) 106

Mathematician here specializing in number theory with a side-order of graph theory. I've only had time to start looking at two of them. First, is the Hadwiger-Nelson/chromatic number of the plane https://en.wikipedia.org/wiki/Hadwiger%E2%80%93Nelson_problem proof. As far as I'm aware (and I could be wrong) the technique it uses here is not in the literature, so this is a genuine construction of a new technique. Second is the Erdos Egyptian fraction bound, and for that one it looks like the techniques are about what I'd expect, but I'm definitely still digesting both of these.

For the quasi-Riemann Hypothesis the striking thing is almost the opposite direction. It looks like the AI used standard complex analytic techniques to get the result. But many mathematicians have often thought for years that those techniques would not likely be strong enough to get this sort of result.

I know less about the Unique Games Conjecture https://en.wikipedia.org/wiki/Unique_games_conjecture but having talked with some of the people there it looks like it took a somewhat standard set of ideas and then combined them with multiple just weird stuff and sort of took a hard left turn at one point for no clear reason and ended up at the result.

It is also worth noting that while many of these have Lean code confirming their correctness (quasi-RH for example) others do not. The three I mentioned above have all also been looked at at this point by human mathematicians who have not found issues; that's likely true for others, but those three I'm aware at least of people doing so. Not all the claims have Lean code though; a bit under half. One of the non-formalized problems also has been withdrawn due to what essentially amounts to a sign error https://github.com/openai/math/blob/main/preprints/Algebraicity-of-Weil-classes-on-split-abelian-eightfolds-September-18-2026/paper.pdf. It is likely others will be withdrawn also by the end, but I'd be surprised if more than 10 are. And even if everything single one without Lean code turned out to be wrong (which seems very unlikely), this would still be an amazing set of math. I commented elsewhere that if a human mathematician had made the quasi-RH result they'd be likely a shoe-in for the Fields Medal, and another mathematician replied saying "delete likely."

Now a more editorial comment: There are legitimate concerns about what this is doing to mathematics. This sort of thing is very cool. But it also is part of a trend that may make it much harder to train young mathematicians or get them to exist at all. If the AIs are limited in how genuinely novel their ideas can be, then we may end up in a situation where we get a massive burst in math over the next few years, and then math stalls out because we don't have enough good really high caliber mathematicians (Not the mathematicians like me, but people like Serre, Tao, Scholze, Clausen,etc.) to come up with deeply new ideas that the AIs can build on.

Comment Re:I hope they realize people hate slop (Score 1) 79

people hate AI slop, and doing all these slop-enabling AI things is just going to erode the trust in the actual market.

This. No good game was ever created from a single prompt. Not to an AI, not a human development team. Every single person writing those damn "look which game my AI built from a single prompt and $x in tokens!" postings is a newbie and almost certainly has never actually shipped a single game.

Game development can be summed up as endless iterations. Every system you build, every visual you create, every sound, text, button, movement, weapon, skill, powerup needs polishing, refinement and changes. It is never, never good enough the first time.

And no, Roblox is not a serious gaming platform.

No it is not, but it IS a very successful platform for a specific subclass of games. If I were in it for the money, I could probably AI refactor some of my earliest games and publish them on Roblox and make a quick buck. If your thing is the kind of games you played on your C64 back in 1793, then Roblox got you covered. Well, if you ignore the scams and money-grabbing.

Comment oh Unity, how you lost your way (Score 1) 79

*sigh* I wish Unity would actually complete some of their dozens half-finished sub-systems instead of constantly chasing the latest trend. It's become really, really annoying. The engine used to be really good. These days, it falls apart if you look at it the wrong way, and at least from the responses to my bug reports it's clear that fixing issues is way down on the agenda.

Maybe this is their attempt to regain grounds in the indie game market - the very market they pushed away by focussing on AAA requirements for years and ignoring the single devs and small teams. If so, I can already tell them that it won't work. We collectively laugh about all the "look what my AI built for me from a single prompt" postings on reddit, where people proudly show off some AI slop that would get them last place at a game jam.

Their AI efforts so far: Tried to build and sell their own AI tools, to a collective yawn from the audience. Built a barely-working MCP server, then deprecated it before it was even finished in order to push a CLI interface to the editor instead, which only works if it's run as the same user as the editor on the same machine. And completely ignores every single advise on workflows and good development practices you can think of.

Unity, I love your engine, I despise your management. Please return to when you were actually good and stop chasing butterflies.

Comment Re:Paywalled AI Marketing (Score 2) 78

It's a regulation forced on them by the EU's AI Act, which neither companies nor users want.

The target audience for this particular rule of the AI Act isn't the producers or users of AI. It is the consumers of content that might be AI generated.

And frankly, letting people know if they see AI content or not is a fair thing to ask.

Comment Re:So what's it good for? (Score 1) 78

And honestly, "not measuring how much a human contributed" is IMHO a massive flaw in general

Fun fact: The EU AI Act considers this. It makes a difference between "AI generated" and "AI modified".

But it's hard to put this into the type of watermarking they are using. Because if you generate a text and then change a few things here and there what you are doing is degrading the watermark. But since the watermark is probabilistic, it's hard to tell if it was a weak watermark from the start or if it was degraded.

Slashdot Top Deals

Arithmetic is being able to count up to twenty without taking off your shoes. -- Mickey Mouse

Working...