Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
1 points by mpweiher
1 points by mpweiher
Is it just me or are local models improving much more quickly than the frontier models?
If so, that would be a very welcome development, with local/private inference becoming available to everyone.
The fact that you get 98.2% of the benchmark performance after going from 16 bits per weight to 1.76 bits per weight really shows you how much space is wasted in these models. Wow.