Why Are Coding Agents So Dumb?

27 points by mtlynch


kingmob

I get that this is more of a rant, and I agree with a lot of it, but some of these are Claude-specific or have plausible reasons.

For example, I don't think harnesses parallelize work automatically because they're more likely to hit session limits, and they consume more tokens in message-passing, coordination, and conflict resolution than when serial.

Likewise, when it comes to model/task suitability, self-reflection on model suitability is non-existent in general, yes, but there are passable alternatives.

In Claude, I instruct it to use models appropriate for the subtask, and it kinda works, but if you use Oh-my-pi, it allows you to specify models for different roles (slow, smol, review, basic tasks, etc), and will use them consistently.

Oh-my-pi also directly supports advisor agents, where a smarter model watches another subagent, and interjects as needed, which is a hybrid solution.