No, local models will not win

13 points by liberty


liberty

For me, local, open-source (not "open weight") models are the only things that would cause me to consider using LLMs. If there isn't an endpoint in which we're not renting our brains instead of using them ourselves, what's the point even for those who are enthusiastic about LLMs?

I disagree with the author about LLMs in general, but I agree with the conclusion. If the one people can rent is better, they're more likely to want to use that. There doesn't appear to be a positive end result for users and programmers without the market otherwise crashing.

ohrv

"win" is a false dichotomy - it's a bit like someone saying "local computers will not win" in the 90s and claiming that servers will always be more efficient and users always prefer faster machines.

chobeat

lot of "business as usual" assumption in this article. No collapse of the AI industry, no collapse of the hardware supply chain, no enshittification, no foreign powers bombing data centers, no regulations or mass retaliation against the tech oligarchs who control the closed systems. The AI industry as it is now is not going to last, one way or another. Consumption patterns have proven completely irrelevant for its development, so if we are going to use local models or not is not really a matter of consumer preferences.

mrunix

The argument that local LLMs will not "win" because they'll always be less powerful than frontier models is like saying that personal computers will not win because mainframes will always be more powerful

strugee

The author doesn't define "efficiency", which is a problem.

Also, this claim:

almost everyone’s revealed preference is to use the strongest available model in their price range

is unsubstantiated and as @dualvariable points out, makes a likely personal experience-based (read: likely heavily biased) assumption about what people are wanting out of these things, on average.

dualvariable

Most people don't need a model that can solve the Riemann hypothesis in order to give them a recipe for the ingredients they have in their fridge. Or even to do a web search and summarize the results.

maduggan

I think this is a classic case of the author assumes their usecase is the same as everyone else. I don't want agents that do more without human intervention. I don't trust agents to do that sort of work unsupervised. My ideal usecase is something I can assign low intensity boring tasks that it can safely grind on in the background and then I check the output of what it produces before it does anything with it.

mdaniel

The author stopped typing after the word "win" since there are a great deal of things one could possibly "win" and it's pretty presumptuous to imply that renting compute is the only way to win all of them. If that were true, no one would have a general purpose computer because the cloud is better, or Amazon Luna or GeForce Now would replace gaming rigs, etc

frontsideair

almost everyone’s revealed preference is to use the strongest available model in their price range

That is true for now, but once we hit diminishing returns it may change. At work, I do not use Fable, or even Opus sometimes, because they have become so similar in terms of capabilities. And providers also have to make the models even more proactive to get over the diminishing returns, which I don't prefer personally. For personal projects. Qwen3.6 27B is enough for most of my needs.

datacenter models are always going to be cheaper

That's a weak reason, I prefer local models because of data privacy mainly. I also like that I will always have access to it, and it works offline. And my GPU is both for gaming and inference, so it's not single purpose.

In the end, as they keep building more datacenters, local models may not "win" but that win condition is a false one. Local models don't have to be better or fully phase out cloud models. I think they are already good and cheap enough, and they'll only get better at closing the gap.