A Severe Misalignment of AI in Mathematics
40 points by jo3_l
40 points by jo3_l
Feels identical to what is happening in software to me.
In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas.
Sounds like what I try to do every day. All that old schpiel we used to say about software being the accumulation of questions and errors isn't nonsense.
The situation seems analogous indeed. However, cynically, I do wonder whether the outcome will differ across the CS and math industry once the dust settles. One could argue that software production is funded in large part for the end product, not the process, whereas mathematics is far more exploratory. More concretely, it seems to me that an agent that outputs unreadable code implementing software that works perfectly from a user perspective is substantially more useful and disruptive than an agent that outputs a unreadable proof of a mathematical result that is formally verified.
Aye, it's hard to deny that software has a function at least somewhat divorced from the process of its creation, although it's still up in the air for me as to how much.
I am oddly optimistic to see this issue so starkly realized in the field of mathematics. If we can come to consensus that "proof factories" are not really helpful, that's at least a grounding point for the rest of knowledge work.
It's also identical to what HAD happened to software cf. The Rise of Worse is Better.
Human slop happens well before AI slop. Now it's only getting worse!
One question I have is whether/how AI "solutions" make it back into future models. I've sort of naively assumed that AI companies basically scrape the web in large swaths, do a bit curating, and train. In a case where an AI solves math problems, I conjecture, that lots of text is created documenting that the discovery happened (only because its new), but very little where mathematicians are spending the time to really document and discuss it in detail where future training runs will have lots of human produced content to ingest. Are Anthropic/Google/OpenAI specifically feeding back these refined gems into future training runs with a "this really matters, but since we took away the thunder of mathematicians online geeking out about it, you should still consider this as as significant as other great discoveries from the past"?
I think your question is adjacent to Model Collapse? https://en.wikipedia.org/wiki/Model_collapse
Not directly, as you're focused on how specific knowledge gets emphasized/deemphasized as a second-order effect of LLMs "taking over" a field, but it seems like a similar issue.
In at least some conditions retraining on e.g. LLM-generated proofs might eventually lead to degeneration, but that doesn't seem to be a foregone conclusion. Maybe other models can be developed to automate curation for future training, even.
It really does seem like the ultimate "best case" scenario is that humans aren't needed anymore. Everything is solving for that. It would likely be best for the other living things on Earth anyway, I guess!
The quote at the top is inspiring but not realistic. Humans have long produced proofs that are difficult to follow and there are many proofs that only a few understand after long and specialized study.
It seems out of place to me to complain about mathematical discoveries made by computers simply because they might be hard to understand by humans.
I even thing it will go in the opposite direction: we will be able to train LLMs to break down mathematical proofs tailored to each of our indvidual levels and ways of understanding, thereby improving mathematical education.