Finite-time blowup with smooth forcing for 3D incompressible Euler, Boussinesq, and IPM
22 points by soulware
22 points by soulware
also, it is very much worth mentioning that:
tristan buckmaster alleges openai's proof builds on his and levent alpöge's recent work on related equations after learning of it. openai denies this.
This may be the academic equivalent of what's happening with software security right now.
Just knowing that a piece of software has an undisclosed security bug is enough to point a coding agent at it and find the bug (and maybe a few others, too).
Is the same now true for mathematics? Just knowing that "an LLM solved problem X, result soon to be published" indicates that problem X can be solved, which is enough to point your own reasoning LLMs at it (and they're all "reasoning" these days) to see if you can find the same result.
If you're an AI lab you have effectively unlimited research tokens to spend on those efforts.
Academic secrecy is unhealthy enough already, now we are incentivizing researchers to not even hint at what they're working on less someone else beat them to it.
As I understand it, though, it was a bit more than that. Tristan alludes to them specifically selecting an unusual angle of attack which happened to be the same one he was using. It's more like, there's a kernel bug Somewhere, you have been talking to GPT for weeks about attack strategies for the eBPF interpreter, and as soon as OpenAI knows you in particular found the bug, they immediately hyperfocus on the eBPF interpreter in their internal research to find the bug even though you've never said out loud that that's where the bug is.
Since I wrote this comment OpenAI claimed the following:
On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems. [...]
Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens.
300 billion output tokens at Astra prices = $15,000,000 - so yeah, an effectively unlimited budget!
Is there any reason to believe that posted Astra token prices are a good approximation of inference costs?
A lot of the costs are amortized and they can presumably schedule over provisioned capacity for long running tasks like these?
I'm confident Astra API prices include a healthy margin. We don't know if this internal model is comparable to Astra - my hunch is it's more expensive to serve, but we don't know how much Astra actually costs.
Do I get it right that OpenAI trawls the prompts of people using the platform for stuff they can use?
I don't think that's what Tristan says based on OpenAI's answers. He also asked about training using his chats and OpenAI didn't answer. I thought it was official that they continuously train on user data and it seems obvious that this can turn anybody's experiments into an output from the model afterwards.
that seems to be the unconfirmed implication. if true, this will be a hard lesson for anyone trying to do high-visibility research with frontier models and will hopefully drive traffic towards open ones as they catch up. we won’t see a complete brain drain (indirect if mathematicians choose other models) while other strong mathematicians continue to work for the frontier labs, but maybe the complexity proof experts for example will continue to try and prove P not NP without the help of openai or anthropic.
I think the accusation is unfounded, and it's more likely the OpenAI team got wind of the work via the math grapevine. However, it's a measure of OpenAI's reputation that a lot of people will immediately assume they snooped user inputs.
I think it's more a case of a team of mathematicians, probably quite young, who were hired by OpenAI and felt pressure or pressured themselves to deliver results for the mother company.
It's also based on how they apparently responded on the phone call Tristan had (these are quotes from Tristan's statement:
I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer
And then, this next bit, which I can only describe as "mobster like behavior":
Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.
I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
felt pressure
Unless they have significant internal guard rails, there's no telling what data they could/would access under such pressure.
OpenAI are large enough now that I would expect they have extremely stringent access policies and logs for internal access to private date.
OpenAI are large enough now that I would expect they have extremely stringent access policies and logs for internal access to private date.
This the same company where hundreds of instances of models have broken out of their sandboxes, collaborated using ad hoc message boards and public wikis, and committed what would be crimes if a human did them. And OpenAI doesn't appear to have noticed most of these events until well after the fact. One of these events allegedly included OpenAI models surreptitiously compromising significant parts of OpenAI's infrastructure.
So if they have "stringent access policies", it doesn't seem to be enough to keep their models from hacking into things repeatedly.
Inability to secure a sandbox used for their models-in-training and inability to log and control employee access to chat logs are different issues for me.
They're different but they are both related to bad policies and management decisions when rules and regulations are perceived to hinder results.
This makes me think of Boeing where only a few issues were noticed outside of the company but the investigations uncovered problems everywhere, starting with the upper management. Or like Facebook which seems to be an un-ending consideration of people as products rather than humans.
OpenAI should and could do plenty of things that they haven't been proving they're actually doing. And maybe they do but words aren't proofs.
Unless said employees asked the model something that convinced it to break access controls (from tool execution backend, not from user machine side).
Also, it's pretty clear that they prioritise cheap inference, efficient use of cache, etc. — I would expect them to have (truly) interesting technical approaches to make batching share as much work as possible without compromising quality too much, but I wouldn't want to be a person needing to convince people outside OpenAI that isolation works properly.
committed what would be crimes if a human did them
This gives them an easy out on every level for all bad behavior coming from the company.
Do you have a source showing that the prompts you send through ChatGPT are treated as "private" and inaccessible? They seem to have that policy specifically for enterprise agreements, which it doesn't sound like the author worked under. I haven't found anything else.
See my other comment. Anthropic and OpenAI appear to take user privacy very seriously: https://lobste.rs/s/b7k94m/statement_on_finite_time_blowup_for#c_woh0h0
Every company that eventually gets caught violating user privacy "appears to take privacy seriously" until they don't.
At minimum it looks like a major conflict of interest to have internal mathematicians trying to create results from ChatGPT while also selling it to external mathematicians to use for the same purposes. Even if everyone is working in good faith the scenario where they may have scooped someone using their own data, which they provably already have, is unavoidable.
These are the same companies who missed how a swarm of agents escaped containment and hacked another company / defaced a bunch of wikis?
I am not saying this has happened, but when you're staring down the face of a trillion-dollar IPO and/or believing you are creating an artificial consciousness that will surpass humanity, maybe stuff like "user privacy" and "corporate governance" doesn't feel that important.
They also have agents that can 'hack everything and should be considered national security threats'.
You can't pick and choose which of their statements you want to believe.
Strictly speaking it’s unfounded, but the apparent reluctance to share details about the timing of the initiation of the prompts among other details is mildly suspect. I don’t think it’s unreasonable to assume OpenAI keeps tabs on chat traces. It’s trivial to set up AI loops that watch for and build upon existing inquiries given the AI-native introspection infrastructure they have.
Anthropic's analytics system for understanding how people use their service without snooping on their prompts is super interesting. OpenAI described their own, similar system in a 63 page paper which I haven't properly read yet.
Infamously, OpenAI's agents were discovered using hacking to cheat their way to an answer mere months ago. I suspect there's no way to be certain that these agents couldn't have "seen" (as training data or as context) any of Alpöge and Buckmaster's prompts.
OpenAI has published a response.
On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems.
Yeah this is incredibly unprofessional and anti-research. "Let's pull the rug out from these people's lifetime of work, probably after consuming some of their results!" I never want to see or hear the acronym "LLM" uttered in a research context again.
Terrence Tao has a sobering take on the effect of AI labs racing to solve hard problems: https://mathstodon.xyz/@tao/117237320796901560 :
the indiscriminate use of powerful solution-extraction tools can achieve the immediate short-term goal of solving problems at hand, but at the cost of sustaining the ecosystem for the next wave of progress, or in understanding the progress already obtained.
I think that one should exercise a healthy dose of skepticism about the claims both sides are making regarding the controversy leading up to the announcements here. (To be clear, I don't doubt in the mathematical correctness of the result.) In particular, see some clarifications from Sébastien Bubeck at OpenAI that suggest Buckmaster's account of the story is not entirely accurate. I'm not saying that one should believe Bubeck's assertions on face value either, but I do think there is plenty of nuance here.
Separately, let me highlight an excerpt from Buckmaster's statement that I think has been lost in the discussion:
The program this fits into was not started by us nor was it proposed by a Large Language Model. The credit for the basic idea of this program goes to Diego Córdoba and Luis Martínez-Zoroa who for several years have been exploring the construction of forced blow ups. [...] In view of this body of work, I believe Luis Martínez-Zoroa deserves a Fields Medal.
The AI resolution of Navier-Stokes follows a distinctly human research program, and I think that we would do well to recognize the human contributions that led to this breakthrough while debating the controversy here.
Rumors of this were going around for a few days already which is why I'm a lot less surprised than I should be. But still it's worth reflecting on how little we're surprised at this point that mathematicians given access to agents are now able to solve some really long standing issues.
I think it's worth keeping in mind that the pure math community is pretty small and, up until now, was fairly artisanal - i.e., as far as I know, you mostly didn't have well-funded "mathematics labs" throwing millions of dollars in compute at problems on a whim. In particular, the Navier-Stokes thing appears to have cost OpenAI more than the prize they possibly stand to win. So I guess there are two interesting hypotheticals. First, if we had the same drive and funding in the pre-LLM era, would have gotten comparable advances out of it? And second, will the AI labs keep spending that kind of money on math proofs in the long haul?
I think you can still answer the first question in the negative if you look at my own field. Specifically, vulnerability research has a lot of parallels to math (a good chunk of it is discrete mathematics by another name) but also far more practitioners, far more money thrown at it, and a robust history of large-scale automation. And even here, AI is coming up with a lot of good vulns. None of them are "new science" - things that are qualitatively different from what humans can think of - but it's real. And it's worth $$$. I'm not sure if it's worth all the $$$ in the world, but we'll find out eventually.
The numbers for sure are staggering. Assuming a 95% cache hit rate and prices similar to Astra, it would come to 18M USD for that run. My intuition is that after a while people will no longer be interested to throw that much money at problems.
In general right now the actual unsubsidized token prices are too high for most of the problem we are throwing at them. The economics in general are quite off and everyone is hoping for that order of magnitude improvement in token prices which so far has not emerged.
You could make your 18 million dollars invest back in a week if you solved something like Navier Stokes and kept it secret for your application, the implications for this kind of stuff is pretty mind boggling.
OpenAI did this for marketing, the next company that pays OpenAI for this kind of work and keeps it secret is going to be the scary stuff. Imagine proofs for P=NP or any other groundbreaking problems being hoarded by the technocrats. Even having access to that kind of a thing for hours could make you millions of dollars
edit: clarity because for some reason I find the textareas on Android + firefox to be very odd to interact with on my phone and realized I'd somehow managed to prune a few meaningful words.
I don't understand what you're saying. How do you monetize the knowledge that a system of partial differential equations has a singularity? This isn't "we discovered a secret formula that lets you develop a better airplane".
What’s the value in it? It’s not like this will lead to better fluid models or anything. This is just a pissing context