Comment Re: Jacob Coxon is a nobody (Score 1) 103
When LLMs have a memory longer than an hour, I will become more concerned.
I'm not sure what an hour memory here means, since context windows are measured in tokens. Even six months ago, LLMs could think for well over an hour before producing an output.
For anyone who is actually very deep in AI *in the real world*, not just heads down in foundation labs - the hyperbole is off the charts on all of this. We are NOWHERE NEAR the exponential yet. These labs can't even get AI to improve their own software reliably, let alone do other things reliably.
This seems odd at multiple levels. First, the people in the foundation labs if anything have more of an idea what is happening than the general population. They are seeing how models are made, what they contribute, and how they go run off the rails. I'm also puzzled by the claim that we're not at the exponential. Two years ago, AI systems could barely do AMC and AIMIE problems, basic high school math contests. Then the AI got to be good enough to do well on Putnam results. Then, early this year, AI got to be good enough to solve multiple open unsolved math problems, including Erdos 1196 https://www.erdosproblems.com/1196 , the Unit Distance Conjecture https://arxiv.org/abs/2605.20695 and many others, and the trends there are just continuing. That's one field, but it is one where the standards are most objective about what is happening. That certainly looks like exponential growth. I'm also puzzled by seeing the labs cannot get the AI to improve themselves is a good standard. From the perspective of a lot of people who are concerned, once that's happening, it is likely too late.
Yann is the most sane of them all because he recognizes we need some massive breakthroughs if we want to achieve real AGI
What you mean here is you agree with Yann. But the point is that lots of people who are as qualified as Yann, disagree. So if, as in your first comment, your primary problem with Coxon is lack of experinece, you should find this situation alarming. Worse, your complaint about Coxon was lack of lab experience, but in your new comment you object to people being too deep in the labs to be objective. This leads to the weird situation where no one is in a position to be relevant to listen to.
. LLM tech has hit its limit, unless someone cracks memory and continuous training.
It is possible that LLM tech has its limits. But people were saying it was hitting its limit 4 years ago, and 3 years ago, and 2 years ago, and a year ago. 4 years ago, people who were saying LLMs were limited would not have predicted they'd be solving unsolved math problems. All those predictions turned out to be wrong. Why should saying it now be more likely? And what if you are wrong here?
Comment Re:These PR campaigns are getting wild (Score 1) 44
Comment Re:Americans love the flyovers. For many man decad (Score 2) 147
Comment Re:Jacob Coxon is a nobody (Score 1) 103
Comment Re:These PR campaigns are getting wild (Score 2) 44
Comment Abusive tactic (Score 4, Interesting) 147
"The flyovers will continue until morale improves," it said.
The original quote they are riffing on is "The beatings will continue until morale improves." Essentially this is a variant of the "make a threat, but in a joking fashion, but really functionally a threat" sort of thing that dangerous and abusive people will do. Trump uses versions of this all the time, and it isn't surprising that Hegseth and his people have picked up on this. I'd call this another warning sign about fascism but given everything else that's going on, including a grossly illegal war in Iran, Trump functionally threatening to undermine the upcoming election, and Trump having removed generals and other higher-ups in the military that he sees as insufficiently loyal (read, more likely to actually follow the law and morality when given an illegal order), and this incident becomes almost part of the rounding error.
Comment Re:what? (Score 1) 128
Comment Re:Sounds about right (Score 1) 128
Comment Re:Read Bubeck's response, not just Bumaster versi (Score 2) 128
If you have people working on a problem for decades, and there are rumors that they have made progress, what is the motivation to independently launch your millions of dollars of AI inference at that same problem?
In a normal academic context, none and moreover this sort of deliberate scooping would be considered incredibly bad behavior. When you have two giant corporations fighting to be able to have results they can showcase to the public and their investors? Then it makes a lot more sense, even as it is incredibly damaging to the academic math culture. This is really toxic behavior and deserves pushback. But it is a mistake to go from there to think that therefore any copying or plagiarism happened.
Comment Read Bubeck's response, not just Bumaster version (Score 5, Informative) 128
I would like to clarify a few things:
1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions.
2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee.
3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)
4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
As I said in that thread, I don't know whose version of events is accurate here, but some of the details due support Bubeck's version. I suspect that a breakdown of communication occurred where both then misinterpreted what the other saying in a more hostile way than it was intended until the conversation then became genuinely hostile. In terms of the math, although I'm a mathematician, differential equations is pretty far from my expertise, so I cannot deeply evaluate how close the methods were. However, it is the case that the both were building on the methods of Cordoba-Martinez-Zoroa. In this context, my being far from this sort of work is relevant, because this was a well known enough approach that even though I'm in a pretty different subfield, I had heard of CMZ's work as an approach and that this was considered a promising approach to Navier-Stokes. Given that, my inclination is that the case that any theft occurred here is very weak.
And since that thread, I've become more convinced, as experts reading both papers point to substantial differences in the details of their approaches. One of the ironies here is that a lot of the anti-OpenAI views are coming very loudly in part due to a general anti-AI attitude. But Bumaster and Alpoge, the people who are claiming to have had work stolen (well, primarily Bumaster, Alpoge is mostly staying out of the fray) were heavily using both Claude and Codex in their work.
Comment Re:Goal (Score 1) 33
Once again: training a model with the reward being "does it solve the task?" without looking at how it solves the task is very, very dangerous.
So, people are able to recognize this, and yet in the other threads here about Anthropic calling for a slow down of AI research, everyone seems convinced that this isn't about the risks really at all.
Comment Re:We are going so fast we need to slow down! (Score 0) 118
Comment Re:We are going so fast we need to slow down! (Score 1) 118
Comment Re:Or, counterpoint, it stole somebody else's work (Score 1) 97
You have never even met a practicing theoretical mathematician let alone participated in a conversation with one, have you? How would you know what constitutes a faux pas in their community, or how they collaborate?
Did you see in the comment where you are replying to where I said I'm a mathematician? My own primary areas of research are number theory and graph theory. It is pretty easy to find who I actually am and verify that yes, I am a mathematician and am in a pure field. So yes, maybe I do know something about the discipline.