Why Are Coding Agents So Dumb?
27 points by mtlynch
27 points by mtlynch
I get that this is more of a rant, and I agree with a lot of it, but some of these are Claude-specific or have plausible reasons.
For example, I don't think harnesses parallelize work automatically because they're more likely to hit session limits, and they consume more tokens in message-passing, coordination, and conflict resolution than when serial.
Likewise, when it comes to model/task suitability, self-reflection on model suitability is non-existent in general, yes, but there are passable alternatives.
In Claude, I instruct it to use models appropriate for the subtask, and it kinda works, but if you use Oh-my-pi, it allows you to specify models for different roles (slow, smol, review, basic tasks, etc), and will use them consistently.
Oh-my-pi also directly supports advisor agents, where a smarter model watches another subagent, and interjects as needed, which is a hybrid solution.
Thanks for reading!
For example, I don't think harnesses parallelize work automatically because they're more likely to hit session limits, and they consume more tokens in message-passing, coordination, and conflict resolution than when serial.
I think this is a plausible explanation, but I still find it unintuitive. Wouldn't you expect the agent eat up more tokens to do 10 tasks in the same session/context than to have one supervisor delegate to 10 tightly-scoped subagents?
But even token efficiency aside, I imagine there are lots of users like me who don't run up against quota limits often and would happily trade tokens for faster execution (in wall time).
Wouldn't you expect the agent eat up more tokens to do 10 tasks in the same session/context than to have one supervisor delegate to 10 tightly-scoped subagents?
Depends on how well the agent manages to provide meaningful context to the subagents, and they don't end up rediscovering the same things the main agent already did. In practice, this seems a bit mixed.
"If I had a human employee tell me they sat idle their whole shift because they wanted my input on some superficial detail, I’d quickly fire them."
I think I found the root of the issue here.
Your wish list for coding agents is essentially 70 years worth of computer science, design and product discipline, coupled with super human intelligence and human level restraint. If we had all of that we'd have solved pretty much every problem in CS. And you want it to be open source.
We're closer to having humanoid robots that can reproduce themselves than that.
Thanks for reading!
Which wish list items do you mean? The models are capable of these things but the agents don't take advantage.
It's both model and agent. The model has a context limit, so until someone figures out a way to efficiently flush and make it long lived, you have to treat every LLM interaction as if with a baby freshly spawned but with a vast compilation of knowledge but that knows nothing about the specific environment except AGENTS.md or whatever shitty crutches we have.
Yeah, that's what hold agents back the most right now
You can ask them to take notes into a file, but that solution come with a lot of drawbacks:
So you end up repeating yourself, again and again, until you get completely mad.
This is a great rant :-)
One thing I'll often do is ask my agent to update the project docs first (or I'll make initial edits). I want to see what this will look like from a user's perspective. Once I'm happy with that, I'll move to plan, and then code.
I agree, I rarely read every last plan detail, but they are probably still a useful exercise for the agent to go through the planning step for itself (like a human).
I suspect frontier labs see harnesses as just an LLM-to-bash adapter. You won't see smart harness features from them, because everything a harness could help with, they hope to RL-train into their next model.
a lot of this cognitive dissonance goes away if you just substitute "tool" or "code generator" for "agent" throughout. anecdotally you will also have a much better time using them, in terms of satisfaction vs frustration.
I think the part about OS level sandboxing is factually wrong? Claude supports that via bubblewrap: https://code.claude.com/docs/en/sandboxing
Which part is wrong?
Sandbox is off by default in Claude Code. Even when you enable Claude's sandbox, my understanding is that it still gives itself read access to your whole system, and you have to enumerate every path it shouldn't be able to read. I'm reading the docs now and having trouble understanding how I'd even configure it to say, "You can only read the current directory."
My point is that Claude Code and Codex are insecure by default, and they make you do a lot of work to restrict their access rather than designing the agents to use the least privileges necessary to do their work.