What Would A Serious AI Product Look Like?
83 points by Corbin
83 points by Corbin
I really like this article. In particular:
There is no reason for a software development or research tool to use first-person language to describe itself. They should not do so. In fact they should not be allowed to do so.
This is fantastic, and in a healthier world this would be made law.
Excellent article, enough ideas to justify several, but I'm not complaining.
I have yet to hear from a single person who has said “yeah, we measured according to your methodology
Me neither, and I continue to wonder why. I feel like even I, with my limited scientific background, could run an experiment with reasonable, non-perfect controls that would be good enough for a corporate environment. I suppose the fact that I haven't means that either everyone else is sharing my reason for not having done so (overwhelming fatigue with the space), or it's just much harder than I am imagining.
Fatigue definitely has something to do with it. Also, many folks have already made up their minds. Suppose you're an AI promoter who's a genuine believer—you're not being paid to promote it or anything like that. From that person's point of view, whatever they've already observed must have looked like this, so what's the point of layering on additional scientific rigor? It will either confirm your existing beliefs, or you will think the experiment must be faulty in some way: maybe it's measuring the wrong thing, or maybe those non-perfect controls actually mattered. But your existing beliefs won't update.
Text editor analogy: a study showed editing text with keyboard+mouse is faster than keyboard alone, even though it feels slower. Except apparently that study was bogus? I dunno. The only thing I do know is I've never met anyone who read any of the research and then changed their behavior because of that: e.g., switched to/away from a modal editor.
(I do think we should do more science! I'm just pessimistic that it will change minds.)
This is a good article, but just getting LLMs to present a list of citations is not enough, because those citations will be cherry-picked. The only way to be sure you are getting a balanced view is to actually do your own research, at which point the LLM becomes useless.
This will obviously vary from person to person and task to task, but the question is: what is the acceptable level of accuracy?
It's never going to be 100% perfect. You could hire me, a bona fide human being with pure intentions, to do your research and provide you with a list of citations. But if I haven't had enough coffee, I might provide a citation that doesn't fully support what I think it does. And even if I was perfect, my source could be wrong.
In terms of rigorously maximizing accuracy, I think the closest we currently get is stuff like LLMs generating Lean proofs. There's still the matter of mathematicians verifying that the proofs prove what we think they prove, of course—I guess the bet there is that humans reading the proof will be easier than writing it from scratch.
The argument in the article is that you should check the work of an LLM. My point is you cannot do that by assuming that the citations provided by the LLM are the relevant ones. Likewise if I was checking the work of a human, I would not assume that what they tell me is the whole picture either.
Eg if I ask an LLM to research whether vaccines cause autism and it cites a big long list of plausible looking articles saying that they do, I shouldn’t just take that at face value just because those articles say what the LLM says they say. Likewise if a human researcher told me the same. You have to be systematic in looking for sources.
IMO LLMs are only useful in the narrow domains for which automated checking of answers is possible - software engineering, maths etc (https://neilmadden.blog/2026/04/14/mythos-and-its-impact-on-security/). For anything else, the work to check is the same as just doing the job yourself in the first place.
I would like to see people who are more positive about LLMs (such as @simonw) discuss the article's conclusion. As I understand it, the conclusion argues that the productivity and value that people perceive in these products are entirely canceled out by the risks and harms, and that the AI labs already know this, and that's why they're not eager to add the suggested features. It would be easy to take the lack of a rebuttal as further proof that he's right.
It's an interesting read, but it's also a mix of:
I get what the author wants and they're interesting requests, but:
Most of the interface parts are already possible. But this is unlikely something the top level providers will ever do. It's not how everyone wants to interact with the system - there will be thousands of ways people will prefer this idea was different than the author's perfect solution. So the best thing the model itself can provide is a generic interface with enough features that the right thing can be built on top of it. Some people will want to analyse a pattern in the picture, some will want sentiment check, some will want code, etc. and I agree some of those use cases should get their smaller specific versions, but you can't satisfy everyone - most people will need the generic version. (And tools built on top)
Getting pi with a few existing plugins and some custom ones will provide you the research/citation tool the author wants. It costs some time and tokens, but nothing stops you on a technical level from getting this tool today. Feed the blog post into an LLM and refine on the plan from there. We have the ability to do it today. The same goes for the widgets - instead of redoing the conversation, ask the model to convert the current results into an interactive website / jupiter notebook / whatever. This is how I mostly use the models in similar situation - instead of just "do the thing", it's "do the thing, now create a skill to do it, now extract all the deterministic parts into scripts that drive the process".
There's a lot of the "deterministic" magic in the post and talk about trusted grounded true answers, like it could be done in any way. Parsing meaning from English is not deterministic or solved. Neither is understanding queries and web text relation. Neither is extracting fragments. Etc. Even if it has strict rules and spreadsheets and citations, you'll never get to a perfect system of extracting quotes - they can completely misrepresent the context. Fictional books can contain fictional quotes of real life people. Or be written in ways where quotes are useless. "Does Burgess glorify violence? Give me citations from Clockwork Orange, no summary or editorialising."
The whole intent and sandbox thing is not solvable in general case. You can create a small set of actions you want to allow, but a general sandbox will never exist. Sometimes I want to edit one specific file, sometimes I want to make direct changes to a filesystem that can wipe anything. Sometimes I want to modify the running system. Maybe I want destructive actions? Even in general, deterministic software we've been struggling with sandboxes for decades and only just started to make them not terrible recently. I hope for some improvements, but they first have to come at the OS level (MacOS doesn't even have a realistic non-deprecated interface for them).
I think the author could learn a lot more about the system and I genuinely hope they stick with it enough to create the research tool they want and share it with the world, because it's needed. But this rant is still a bit early.