Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things

8 points by Yogthos


nelson

There's a fascinating discussion on Reddit right now about whether it is overthinking or this is just how a model can do more complex tasks better. It reminds me of how Claude's new model was using way more tokens to do tasks a few months ago, but definitely worked better. Example threads: here and here.