| HN Mirror

Y	Hacker News new \| ask \| show \| jobs

by nisten 357 days ago

If you want to have an opinion on it,

just install lmstudio and run the q8_0 version of it i.e. here https://huggingface.co/bartowski/Qwen_Qwen3-4B-Instruct-2507....

you can even run it on a 4gb raspberry pi Qwen_Qwen3-4B-Instruct-2507-Q4_K_L.gguf https://lmstudio.ai/

Keep in mind if you run it at the full 262144 tokens of context youll need ~65gb of ram.

Anyway if you're on mac you can search for "qwen3 4b 2507 mlx 4bit" and run the mlx version which is often faster on m chips. Crazy impressive what you get from a 2gb file in my opinion.

It's pretty good for summaries etc, can even make simple index.html sites if you're teaching students but it can't really vibecode in my opinion. However for local automation tasks like summarizing your emails, or home automation or whatever it is excellent.

It's crazy that we're at this point now.

3 comments

esafak 357 days ago

Thank you. To spare Mac readers time:

mlx 4bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

mlx 5bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

mlx 6bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

mlx 8bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

edit: corrected the 4b link

link

belter 357 days ago

This comment saved 3 tons of CO2

link

ckcheng 357 days ago

Did you mean mlx 4bit:

https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

link

magnat 357 days ago

> if you run it at the full 262144 tokens of context youll need ~65gb of ram

What is the relationship between context size and RAM required? Isn't the size of RAM related only to number of parameters and quantization?

link

Gracana 357 days ago

The context cache (or KV cache) is where intermediate results are stored. One for each output token. Its size depends on the model architecture and dimensions.

KV cache size = 2 * batch_size * context_len * num_key_value_heads * head_dim * num_layers * element_size. The "2" is for the two parts, key and value. Element size is the precision in bytes. This model uses grouped query attention, which reduces num_key_value_heads compared to a multi head attention (MHA) model.

With batch size 1 (for low-latency single-user inference), 32k context (recommended in the model card), fp16 precision:

2 * 1 * 32768 * 8 * 128 * 36 * 2 = 4.5GiB.

I think, anyway. It's hard to keep up with this stuff. :)

link

wkat4242 356 days ago

Yes but you can quantise the KV cache too just like you can the weights.

link

hnuser123456 357 days ago

A 24GB GPU can run a ~30b parameter model at 4bit quantization at about 8k-12k context length before every GB of VRAM is occupied.

link

iamnotagenius 356 days ago

Not quite true. Depends on number of KV heads. GLM4 32b at IQ4 quant and Q8 context can run full context with only 20GiB VRAM.