What is Codemode
9 points by sp6370
9 points by sp6370
Are Pi's Codemode and Claude's dynamic workflows essentially the same?
Note that Codemode is by default only enabled in Pi when MCP is enabled
As somebody who does not use Pi I am really curious what led to this decision. Maki enables codemode unconditionally and for all tools, and it didn't cause any obvious issues, besides some models over-using it when it's not necessary.
As somebody who does not use Pi I am really curious what led to this decision.
Combination of conservatism and basic benchmarking across different LLMs. Codemode works very poorly for some LLMs due to how tool calls are encoded. Anthropic special cases toplevel string arguments in antml to be mostly escaping free, and OpenAI enables something similar with freeform tools. So codemode itself can be escaping free on those models. But then the tools within it, lose that optimization because now you're doing JSON escaping.
Some smaller models really suffer from this but they still reach for codemode too much after a while. This will need more investigation before we subject people to that as a default experience.
Codemode also is quite beefy in the system prompt itself, for the agent to make sense of it.
But then the tools within it, lose that optimization because now you're doing JSON escaping.
Ah I see. Maki's codemode is a Python sandbox and tool args are just regular python kwargs. I wonder if that makes a difference. I don't think any real benchmarking has been done
It probably makes a huge difference because the models that are RLed on Codemode were trained on JavaScript. So you are already off the beaten path. But for the best in class models, it won't be very noticable.
Combination of conservatism and basic benchmarking across different LLMs.
This is something I have been curious about for quite some time: How are you doing benchmarking in Pi? How much effort is spent to co-evolve Pi with different models?
From the release notes I am aware that you sometimes optimize the integration of OpenAI or Anthropic models, would Chinese LLMs benefit from this as well, or are they very similar from a harness perspective?
How are you doing benchmarking in Pi?
Badly.
How much effort is spent to co-evolve Pi with different models?
More than you think. But at lot of this right now is vibes based and looking at user reports. A big issue with all of this is that almost anything works with SOTA models (except when it overfits) but with the Chinese open weight models you have all kinds of crazy regressions.
There are some harnesses that go really far patching around some of the failure conditions. For instance smaller Qwen models go into loops like crazy and some harnesses spend a lot of effort to detect loops and break out of them. We don't do any of this because our expectation is that models fix this and we don't want to carry nonsense like this in the core.
would Chinese LLMs benefit from this as well, or are they very similar from a harness perspective?
Not all Chinese models are the same here but outside of benchmarks there is a huge difference between US SOTA models and Chinese open weight models still. But we do not try to benchmaxx or fine tune for a particular model. I know that some are, but it's very hard to evaluate that this actually works well for all situations.
Fascinating. Pi works quite well with Qwen3.8 (27B and Flash Next) and with DeepSeek V4 Flash 0731. Flash Next is very happy with codemode. If those models have problems, I haven't seen them.