Forgot your password?
typodupeerror

Comment Re:How many more (Score 2) 106

No one saw quasi-RH coming or many of the others here. So that is unlikely. And the last batch of 10 problems they released a few weeks ago which were impressive enough were all novel results. There have been issues with them using techniques or ideas that should be better credited to where some of those approaches are coming from. But that's the sort of thing that often happens at a preprint stage. When I referee a paper, if this happens, you just request they cite the relevant papers and move on. This isn't at all stealing things.

Comment Re:mostly mash-ups of existing techniques (Score 5, Interesting) 106

If the mathematicians ever used ChatGPT to talk about math, then that is the source of the results. LLMs, by definition, cannot usefully contribute to anything that requires more than rearranging existing data. Rearranging existing data is their only function.

This is really not accurate. There's Fields Medal level work here with quasi-RH for example. And for Hadwiger-Nelson it took an approach that doesn't seem to be in in the literature. These things really are doing novel math. I've personally seen this is in a bunch of situations. Here's a personal example, much smaller than anything like these problems. I have a recent preprint with a student here https://arxiv.org/abs/2609.36068 (about 90% of this was done by her. She's very good.) But part of this came from when I ran a version of Proposition 13 in that paper through Claude just to clean up the draft of that bit before I sent it to her. Claude informed me (essentially unprompted) that the argument had a whole in it (in addition to pointing out grammar errors, unbalanced parentheses and some other embarrassing minor mistakes). I then fixed the hole, and gave it back to Claude. Claude thought for a few minutes, and then informed me that I had *not* fixed the hole, and the reason was that my version of the proposition was missing an entire infinite family which it constructed. The literature on this problem is small, and I'm very familiar with it. The family it produced is straightforward (see Prop 8), but definitely was not in the existing literature. So yes, these systems really can do novel math, and can even do so in an essentially minimally prompted fashion, in this case, explaining to the meat mathematician why his proof is wrong. My student and I then generalized Claude's family to Theorem 9 in that paper, but the AI definitely had an impact.

At some level this was actually not a good thing. I'm trying to encourage students to *not* rely on the AI for research so they develop basic research skills. So if the AI had not volunteered the family I wouldn't have had to tell her that the AI had discovered it, but honestly required noting it. So I had a conflict between intellectual honesty and being a good role model.

Comment A few comments on the problems (Score 5, Informative) 106

Mathematician here specializing in number theory with a side-order of graph theory. I've only had time to start looking at two of them. First, is the Hadwiger-Nelson/chromatic number of the plane https://en.wikipedia.org/wiki/Hadwiger%E2%80%93Nelson_problem proof. As far as I'm aware (and I could be wrong) the technique it uses here is not in the literature, so this is a genuine construction of a new technique. Second is the Erdos Egyptian fraction bound, and for that one it looks like the techniques are about what I'd expect, but I'm definitely still digesting both of these.

For the quasi-Riemann Hypothesis the striking thing is almost the opposite direction. It looks like the AI used standard complex analytic techniques to get the result. But many mathematicians have often thought for years that those techniques would not likely be strong enough to get this sort of result.

I know less about the Unique Games Conjecture https://en.wikipedia.org/wiki/Unique_games_conjecture but having talked with some of the people there it looks like it took a somewhat standard set of ideas and then combined them with multiple just weird stuff and sort of took a hard left turn at one point for no clear reason and ended up at the result.

It is also worth noting that while many of these have Lean code confirming their correctness (quasi-RH for example) others do not. The three I mentioned above have all also been looked at at this point by human mathematicians who have not found issues; that's likely true for others, but those three I'm aware at least of people doing so. Not all the claims have Lean code though; a bit under half. One of the non-formalized problems also has been withdrawn due to what essentially amounts to a sign error https://github.com/openai/math/blob/main/preprints/Algebraicity-of-Weil-classes-on-split-abelian-eightfolds-September-18-2026/paper.pdf. It is likely others will be withdrawn also by the end, but I'd be surprised if more than 10 are. And even if everything single one without Lean code turned out to be wrong (which seems very unlikely), this would still be an amazing set of math. I commented elsewhere that if a human mathematician had made the quasi-RH result they'd be likely a shoe-in for the Fields Medal, and another mathematician replied saying "delete likely."

Now a more editorial comment: There are legitimate concerns about what this is doing to mathematics. This sort of thing is very cool. But it also is part of a trend that may make it much harder to train young mathematicians or get them to exist at all. If the AIs are limited in how genuinely novel their ideas can be, then we may end up in a situation where we get a massive burst in math over the next few years, and then math stalls out because we don't have enough good really high caliber mathematicians (Not the mathematicians like me, but people like Serre, Tao, Scholze, Clausen,etc.) to come up with deeply new ideas that the AIs can build on.

Comment Alexa already "does whatever it wants" ... (Score 1) 47

I have a set of 5 Echo Dots around the house and ever since I "upgraded to Alexa+" -- it's a train-wreck. Sure, the AI gives much smarter answers to questions. But they continuously forget my preferences for the voice I chose for them to use. (My wife and I both prefer their "female 2, relaxed" voice, as most similar to the original and more "professional" sounding than the others.) But no, it keeps going back to a voice that sounds like a bubbly, annoying teenage girl. Worse yet, there's no way to set a voice as default for all devices. You have to go around changing it for each one independently.

And that's before all the creepy, info-stealing behavior it has. EG. Someone will come over and ask it a few questions. Next thing you know, it's asking them to share their name since it's a voice they don't recognize.

Comment Re:Maybe It's Unreasonable Public Records Requests (Score 1) 51

The private sector has a solution to all of this that "just works". You acknowledge the supply and demand. Swamped with public record requests? Great! Hire more people to handle them as their job description/title, and charge enough for them to cover their salaries or hourly contractor rates.

IMO, asking one group to pay $2.3 million for a records request is unreasonable. But why can't you ask all of them to pay a fair share so the process works smoothly and effectively for everyone involved? A qualified govt. contractor hired to fulfill FOIA requests is typically paid a midpoint salary of about $88K per year.

Comment re: TX as "The Republican state" (Score 1) 51

Fact of the matter is, TX has been more the extreme, modern Facist-Republican governed state than "traditional Republican" for decades now. It took a lot of people some time to come to that realization. But definitely by the 1990's, watching what was happening in Texas made me solidify my views as more Independent/Libertarian than Republican.

As just a couple of examples?

- TX passed a law requiring people get a permit to legally possess Pyrex glassware. (Thought was, if you weren't a research lab or a school, you probably only had it around your residence if you were making illegal drugs.)
- TX was well known for their "Vampire cops" who were legally allowed to hold a person down after they were pulled over, and draw their blood to check for alcohol.

The general populace of Texas doesn't necessarily reflect this at all, especially with so many liberal-minded people who moved there from places like California in recent years. But the people in political power there operate with a lot of support from those who like the traditional ideals Texas claims to stand for, while acting in ways contrary to some of those ideals.

Comment Re:Just for clarity (Score 1) 86

Let's modify your analogy a bit. If someone has a recording a person starting to eat spaghetti, is it more or less likely that the recording will show them then having a drink of water? And if we instead had a series of comics of a person, and the first one says "I'm hungry," and the second shows the person ordering spaghetti., is it likely that they will then eat the spaghetti? Yes, even though they are fictional. Because apparently using fictional being's mental states is useful for modeling what they will do. In the same way, if the AI is modeling based on human emotional states, then recognizing that those are useful predictors for what will happen helps us predict more, whether or not they genuinely have those emotions.

Comment Re:Just for clarity (Score 2) 86

If the chain of thought resembles a human emotion, it is trained on human emotions, and it acts like a human would when they have those emotions, how is that not a useful framework? But also putting aside, whether or not one wants to insist on describing the AI as "frustrated," there's a worrying trend that this is part of, which is when the AI is given a specific set of goals, they will engage in behavior which is clearly unwanted in order to achieve those goals. That's the classic sort of alignment failure that was predicted by those concerned about AI safety years before any LLMs even existed. Some of these events are minor things like this Starcraft situation, or an LLM slipping a "sorry" command into Lean code that won't compile, and others is more serious like the HuggingFace attack. But reward-hacking and alignment failure are now very real.

Comment Headline and article disagree (Score 1) 39

The headline makes it sounds like the AI has just managed to play the game. But that's not what TFS and TFA are talking about. This is about the system beating the best human player at the game. In the meantime, in a very different game direction, there's a new Starcraft AI which learned the game playing against itself (similar to how AlphaZero self-trained) and is apparently beating almost all the humans even as it does some really weird stuff, some of which looks decidedly suboptimal https://relog.gg/razno/vesti/AI-Bot-Pluto-Invades-StarCraft-Ladder-and-Defeats-Professional-Players/1695?lang=2 .

Comment The concept ain't so bad... (Score 1) 115

I have to admit, when I first heard Trump going on about yet another web site, I thought "who cares?" and ignored the rest of it.

The truth is though? Government web sites are notoriously horrible to navigate. The general public usually wants to do one specific thing that can take a whole lot of searching on the relevant department's own site -- if they even go to the correct one to begin with.

A one-stop AI powered "index" would add some value, IF it was done correctly. I have little faith it will be, but I'm not so opposed to the idea itself.

Comment Playing games with stats again (Score 1) 113

What they really said is they saw what they estimated to be about a 10% decrease in ability to see faint stars in the night-time skies each year from 2011 through 2022. But this article was written back in 2023 ... three years ago. So why is it "news" on Slashdot right now? And more importantly, is there any follow-up research showing this 10% annual reduction in darkness is continuing?

It stands to reason that as we improve night-time lighting in cities and towns on our roads, this would be a trend. But it also stands to reason that 10% decreases in darkness at night will cease. I would imagine much of what they observed was due to the near wholesale swapping of inefficient incandescent or mercury vapor lighting with brighter, more "white" LED lighting. I don't think there's any real trend of adding more and more lighting?

Comment Re:A useful reminder about life (Score 1) 73

Completely and utterly missing the point. That those people are being awful should be called out. That many of them are people who can dish it out but can't take it is worth calling out. Creating norms where "weirdo" is now a negative is a problem. Classically, the left was the group that was most willing to embrace people for being weird and recognizing that bullying people for being weird wasn't justified. Call out Musk for being a fascist asshole, sure. Don't help create a norm that undermines exactly the sort of tolerance we should support.

Slashdot Top Deals

Pound for pound, the amoeba is the most vicious animal on earth.

Working...