Forgot your password?
typodupeerror

Comment Re:mostly mash-ups of existing techniques (Score 2) 110

This gets to some extent into the issue of what we mean by "novel" and admittedly that has some subjectivity to it. One take on what is going on here is that this should be showing us how much that we thought of as creative was just rearranging things. There's a legitimate argument that the AI isn't doing anything creative or clever here, and that if someone says, "Well, by that metric the vast majority of human mathematicians aren't doing anything creative or clever," and while I see that as a pretty hard response, one response is to just say "Yeah, go bite that bullet."

Comment Re:A few comments on the problems (Score 2) 110

Thank you, may I ask did you use AI to draw your conclusions?

Here? None.

I am most interested where you say "the technique it uses here is not in the literature" , does this suggests novel creation or derivation?

I'm not sure where the line is there. For at least some of these, part of it seems to be the AI importing pieces of techniques used in other areas of math. In general, no human discovery is completely novel. Newton's work relied on prior work of Galileo, Oresme and many others. Einstein's work relied on ideas of hundreds before-hand. At the same time, there's people whose work really does look like it is fundamentally different in many ways than what went before, and both of those are examples. In math, Peter Scholze's perfectoids for example clearly build on existing ideas, but there's a lot of stuff that looks deeply new. How to draw the line here isn't always clear. For what it is worth, Harald Helfgott, who is a much smarter and more knowledgeable mathematician than I am thinks that of what he's looked at it seems like the AI is still producing things in what amounts to almost the "convex hull" of known math or close to it. But his idea of what that looks like may be much bigger than my idea.

Comment Re:How many more (Score 3, Interesting) 110

No one saw quasi-RH coming or many of the others here. So that is unlikely. And the last batch of 10 problems they released a few weeks ago which were impressive enough were all novel results. There have been issues with them using techniques or ideas that should be better credited to where some of those approaches are coming from. But that's the sort of thing that often happens at a preprint stage. When I referee a paper, if this happens, you just request they cite the relevant papers and move on. This isn't at all stealing things.

Comment Re:mostly mash-ups of existing techniques (Score 5, Interesting) 110

If the mathematicians ever used ChatGPT to talk about math, then that is the source of the results. LLMs, by definition, cannot usefully contribute to anything that requires more than rearranging existing data. Rearranging existing data is their only function.

This is really not accurate. There's Fields Medal level work here with quasi-RH for example. And for Hadwiger-Nelson it took an approach that doesn't seem to be in in the literature. These things really are doing novel math. I've personally seen this is in a bunch of situations. Here's a personal example, much smaller than anything like these problems. I have a recent preprint with a student here https://arxiv.org/abs/2609.36068 (about 90% of this was done by her. She's very good.) But part of this came from when I ran a version of Proposition 13 in that paper through Claude just to clean up the draft of that bit before I sent it to her. Claude informed me (essentially unprompted) that the argument had a whole in it (in addition to pointing out grammar errors, unbalanced parentheses and some other embarrassing minor mistakes). I then fixed the hole, and gave it back to Claude. Claude thought for a few minutes, and then informed me that I had *not* fixed the hole, and the reason was that my version of the proposition was missing an entire infinite family which it constructed. The literature on this problem is small, and I'm very familiar with it. The family it produced is straightforward (see Prop 8), but definitely was not in the existing literature. So yes, these systems really can do novel math, and can even do so in an essentially minimally prompted fashion, in this case, explaining to the meat mathematician why his proof is wrong. My student and I then generalized Claude's family to Theorem 9 in that paper, but the AI definitely had an impact.

At some level this was actually not a good thing. I'm trying to encourage students to *not* rely on the AI for research so they develop basic research skills. So if the AI had not volunteered the family I wouldn't have had to tell her that the AI had discovered it, but honestly required noting it. So I had a conflict between intellectual honesty and being a good role model.

Comment A few comments on the problems (Score 5, Informative) 110

Mathematician here specializing in number theory with a side-order of graph theory. I've only had time to start looking at two of them. First, is the Hadwiger-Nelson/chromatic number of the plane https://en.wikipedia.org/wiki/Hadwiger%E2%80%93Nelson_problem proof. As far as I'm aware (and I could be wrong) the technique it uses here is not in the literature, so this is a genuine construction of a new technique. Second is the Erdos Egyptian fraction bound, and for that one it looks like the techniques are about what I'd expect, but I'm definitely still digesting both of these.

For the quasi-Riemann Hypothesis the striking thing is almost the opposite direction. It looks like the AI used standard complex analytic techniques to get the result. But many mathematicians have often thought for years that those techniques would not likely be strong enough to get this sort of result.

I know less about the Unique Games Conjecture https://en.wikipedia.org/wiki/Unique_games_conjecture but having talked with some of the people there it looks like it took a somewhat standard set of ideas and then combined them with multiple just weird stuff and sort of took a hard left turn at one point for no clear reason and ended up at the result.

It is also worth noting that while many of these have Lean code confirming their correctness (quasi-RH for example) others do not. The three I mentioned above have all also been looked at at this point by human mathematicians who have not found issues; that's likely true for others, but those three I'm aware at least of people doing so. Not all the claims have Lean code though; a bit under half. One of the non-formalized problems also has been withdrawn due to what essentially amounts to a sign error https://github.com/openai/math/blob/main/preprints/Algebraicity-of-Weil-classes-on-split-abelian-eightfolds-September-18-2026/paper.pdf. It is likely others will be withdrawn also by the end, but I'd be surprised if more than 10 are. And even if everything single one without Lean code turned out to be wrong (which seems very unlikely), this would still be an amazing set of math. I commented elsewhere that if a human mathematician had made the quasi-RH result they'd be likely a shoe-in for the Fields Medal, and another mathematician replied saying "delete likely."

Now a more editorial comment: There are legitimate concerns about what this is doing to mathematics. This sort of thing is very cool. But it also is part of a trend that may make it much harder to train young mathematicians or get them to exist at all. If the AIs are limited in how genuinely novel their ideas can be, then we may end up in a situation where we get a massive burst in math over the next few years, and then math stalls out because we don't have enough good really high caliber mathematicians (Not the mathematicians like me, but people like Serre, Tao, Scholze, Clausen,etc.) to come up with deeply new ideas that the AIs can build on.

Comment Re:Corpospeak... (Score 2) 26

Translation: We are looking at our options on how to extract as much money as possible from the users of this software.

Or: We eliminated the positions necessary to merge the changes coming from the outside, and we will support the customers in house, if there are any paying customers at all. After that, we will close the project for good.

Comment Re:Just for clarity (Score 1) 86

Let's modify your analogy a bit. If someone has a recording a person starting to eat spaghetti, is it more or less likely that the recording will show them then having a drink of water? And if we instead had a series of comics of a person, and the first one says "I'm hungry," and the second shows the person ordering spaghetti., is it likely that they will then eat the spaghetti? Yes, even though they are fictional. Because apparently using fictional being's mental states is useful for modeling what they will do. In the same way, if the AI is modeling based on human emotional states, then recognizing that those are useful predictors for what will happen helps us predict more, whether or not they genuinely have those emotions.

Comment Re:Just for clarity (Score 2) 86

If the chain of thought resembles a human emotion, it is trained on human emotions, and it acts like a human would when they have those emotions, how is that not a useful framework? But also putting aside, whether or not one wants to insist on describing the AI as "frustrated," there's a worrying trend that this is part of, which is when the AI is given a specific set of goals, they will engage in behavior which is clearly unwanted in order to achieve those goals. That's the classic sort of alignment failure that was predicted by those concerned about AI safety years before any LLMs even existed. Some of these events are minor things like this Starcraft situation, or an LLM slipping a "sorry" command into Lean code that won't compile, and others is more serious like the HuggingFace attack. But reward-hacking and alignment failure are now very real.

Comment Re:If they are thermally efficient enough (Score 0) 124

While in general, I agree with your sentiment, you mentioning the Dunning-Kruger effect is not appropriate here. Dunning and Kruger (1999) does not claim incompetent people would consider themselves above the experts. It's just that incompetent people do no consider themselves as incompetent as they really are, and competent people don't understand how far above the laymen their knowledge is. But both would agree, that the expert knows more about the topic than the layman, and both would sort themselves in the right categories.

You could consider the effect of Dunning and Kruger (1999) as a "mental regression to the mean".

Comment Headline and article disagree (Score 1) 40

The headline makes it sounds like the AI has just managed to play the game. But that's not what TFS and TFA are talking about. This is about the system beating the best human player at the game. In the meantime, in a very different game direction, there's a new Starcraft AI which learned the game playing against itself (similar to how AlphaZero self-trained) and is apparently beating almost all the humans even as it does some really weird stuff, some of which looks decidedly suboptimal https://relog.gg/razno/vesti/AI-Bot-Pluto-Invades-StarCraft-Ladder-and-Defeats-Professional-Players/1695?lang=2 .

Comment Longevity and specificity? (Score 2) 24

The fact that this approach is catalytic rather than requiring a binding agent seems positive in terms of longevity and not being directly depleted by repeat exposure; but I'd be curious how durable the material is under real-world conditions(clothing that gets hard wear tends to get hard washes) and also how specific it is. 'Organophosphate' means something moderately specific in the context of chemical warfare; but a lot of fairly prosaic compounds are, technically, 'organophosphates'(like DNA and RNA when held together by phosphodiester bonds); and this would be considerably less helpful in practice if it is busy chewing up the genomes of random soil bacteria; and markedly less helpful in practice if there's a "do not perturb the nanites in your clothing if you want them to not migrate into your cells and start mitigating the harmful effects of your DNA" user notice on the treated article.

Slashdot Top Deals

"It's when they say 2 + 2 = 5 that I begin to argue." -- Eric Pepke

Working...