Hacker News new | ask | show | jobs
by geye1234 33 days ago
Curious to hear if anyone has tried running the 2-bit or 3-bit quantization of this. With a bit of investment I may just be able to swing it locally. I already have 96GB VRAM, so with 192GB RAM, which seems to be the most one can find these days with a 4-slot motherboard, I may be in with a shot. Yes, it'd be slow, but I could give it overnight jobs. But I don't know if running at such a low quantization would make it hallucinate with only a small context.

Qwen and Gemma are great, but they need babysitting every 30 mins, which is quite a cognitive load.

1 comments

2 and 3 bit quants are often closer to gibberish than hallucination, and that'll happen regardless of context.

I shouldn't claim too much, I haven't tried GLM5.2 at 2/3 bit quantization, but if I were a betting man I'd put money on "useless even as a chatbot"