Why i'm still bearish on LLMs after Navier-Stokes
44 points by andyc
44 points by andyc
I find these points interesting
the labor costs of rigorous specification can greatly exceed that of direct implementation of an informal specification. the hardware engineering world presents a great case study on this
That tracks with my experience
navier-stokes and statements in pure mathematics like it are the absolute best case scenario for agentic work against rigorous specification. the theorem statement itself is already a rigorous specification. it has undergone decades of auditing by the mathematical community
Basically, asking the question rigorously can be as much work as answering the question. It's sort of a psychological bias that we often don't recognize that
It is at the root of:
COBOL: "business people can use this English-like language to specify their problems easily, no more programmers"
AppleScript: "English-like almost natural language can be used to write programs manipulating other programs"
Inform: "English-like language can be used to write interactive fiction like Zork (this one is relatively successful)"
Specification languages: "all you have to do is declare the specifications in this almost-English language and you will have a program" (actual book title from 1982: Application Development Without Programmers)
UML diagramming tools: "draw the diagrams and they will be turned into code"
Offshoring: "Send an email to folks who speak English in a low-cost-of-living country and your ideas will be turned into code"
I see a lot of people using LLMs as offshore programming teams with shorter turn-around times and exactly the same "No, not like that!" conversations repeated indefinitiely.
People forget, but SQL was originally intended for business people. That lasted like 5 minutes, lol.
I didn't know that, and I'm interested in that kind of thing! Can you share a link you think covers that history well?
When Ray and I were designing Sequel in 1974, we thought that the predominant use of the language would be for ad-hoc queries by planners and other professionals whose domain of expertise was not primarily database management. We wanted the language to be simple enough that ordinary people could ‘‘walk up and use it’’ with a minimum of training. Over the years, I have been surprised to see that SQL is more frequently used by trained database specialists to implement repetitive transactions such as bank deposits, credit card purchases, and online auctions.
How could you even doubt the work done when it just received a (just-made-up-by-a-trump-backer-meat-proxy-crypto-billionnaire-scammer) "prize" ?
One thing I rarely see mentioned to relativise the impressiveness of these maths achievements, is that openai and anthropic employ a lot of excellent mathematicians. Some just got out of uni, some have more experience. This makes the big picture shift from "autonomous AI solves anything" to "top-level mathematicians with near unlimited material means and good pay are able to solve maths problems with a cool new tool". I don't think there nothing to be seen here, but this is certainly nowhere near "navier-stokes has been solved by AI!" which I've read here and there.
Anecdotally, my colleague doing computational fluid dynamics for biophysical simulations (ie, users of Navier-Stoke, dirty engineers, not proper mathematicians) tells me this discovery is of absolutely no impact for their research.
One thing I rarely see mentioned to relativise the impressiveness of these maths achievements, is that openai and anthropic employ a lot of excellent mathematicians.
This is an important point. It sits at the confluence of hype-driven marketing, pre-print culture and general intransparency. Their results do not go through peer review, published as either a pre-print, a press release or a corporate blog post. They are dumping a lot of stuff on the scientific community and are protected by strategically gullible reporting and Brandolini's law. I don't like that these labs, about to revolutionize science by their own description, ignore established processes. Science should be more open not less.
I mean, these labs are desperate to point to any win and spread as much fear and uncertainty as possible because that's the only way to justify the valuations they are targeting for their IPOs.
I wouldn't qualify this as bearish. All things considered LLMs are and will be a groundbreaking tool. But fair points