What GLM-5.3 Flash running on Chinese hardware actually means
4 points by martinald
4 points by martinald
That's a cool summary. One bit that I thought could use more context was: "As models get larger, you have to split them ..." And yes, that's true, but also the latest model which people loved so much from zai is 320B parameters, compared to fable's 6T. It seems like anthropic and oai sometimes throw more parameters at a problem just because they can instead of working harder within some limits - it seems we don't need that much bigger models to create something great.
Yes: the workaround for worse hardware might be 5.3-Flash and models like it. He's right that there's always going to be a relative efficiency difference, but if fuel is expensive you'd rather be selling a moped than an SUV, so to speak.
A few smaller models now can apparenly tackle implementation work pretty well according to both users and benches (5.3-Flash, DsV4-Flash-0731, Luna). Three of them suggest it's not a fluke. Surprised I haven't seen more reaction to that so far.