Hacker News new | ask | show | jobs
by Tenoke 25 days ago
I really want a 128gb+ machine but it's brutal to be at only 256 GB/s for $4k (especially with the drawbacks of both ARM and AMD).

I fear that by the time the RTX Spark comes out it'd have to be $6k, and by the time a 128gb or more machine with 700+ GB/s comes out it'd be at $10k, way out of most consumers' hands.

Edit: capitalized gb/s to GB/s.

4 comments

A Mac Studio is a much better buy in terms of memory bandwidth, but impossible to buy in a 128 GB configuration. Honestly there aren’t great options right now and it’s probably better to wait for the market to be less insane.
I looked for one and it's impossible to find, let alone at a reasonable price + it does suffer from being harder to train/use less common models and workflows (e.g. arbitrary comfyui ones). Spark at least doesnt have that drawback, while AMD has both drawbacks.

Waiting for the market to be less insane is somewhat akin to waiting for the s&p500 to drop a decent amount so you can buy in.

No, equities naturally trend up with economic growth. RAM is up because of a supply shock, as new capacity comes on prices will drop, it’s a commodity.
Yes, in the very long term, but in the medium term where the listed capacities even matter we are not close to that. Really, the cost per gb of vram on flagships hasnt went down since 1080ti and thats not accouting for the recent increases which will likely last for years.
There are some minor exceptions but agreed otherwise. Manufacturers have been excessively using VRAM limitations as a price point distinguisher.
But it could get worse for two years yet.
If it flys, floats or FLOPs it is better to rent.
Asus rog 13 128gb is still sub $3k last I checked and ticks the boxes. If you get tired of AI it's a kick ass tablet except the weight and can run AAA games for now decently enough.
Is that a notebook? As the thermals don't really compare for a notebook and these high performance minipcs.
> Waiting for the market to be less insane is somewhat akin to waiting for the s&p500 to drop a decent amount so you can buy in.

I said less insane, not sane. If prices go down 20% but are still up 150% over a few years ago, that would still be an improvement from now.

Or it could work the other way: if new hardware comes out in 2027 such that the tokens/$ ratio works out better, that would also be less insane.

> Waiting for the market to be less insane is somewhat akin to waiting for the s&p500 to drop a decent amount so you can buy in.

lol this is so wrong it's funny - equities go up in price, commodity goods go down in price. the two markets are literally diametrically opposed.

I'd have a better portfolio right now if I invested in RAM instead of equities.
RAM is more like agricultural products (with short shelf-life) than commodities like fossil fuels, mineral ores, etc. You can manage an inventory or speculate on production, but you cannot really hold a "portfolio" of it in any sensible way.

So, you should get into RAM futures if you believe this is more than a transient arbitrage sort of situation. All extant RAM will become obsolete as the demand shifts to newer, fancier versions.

>but you cannot really hold a "portfolio" of it in any sensible way.

xAI effectively did and lucked out to cover their losses and more with it.

> All extant RAM will become obsolete

Yes but on what timescale? Replacing the > 5 year old sticks in my laptop would currently cost well over half what the machine ran me when it was brand new.

The Mac Studio M3 Ultra with 96GB of RAM with 1Tb SSD costs $6799, but has triple the memory bandwidth [1]. It's almost twice the price of Strix Halo and when the M5 Ultras come out it will be interesting to see how much they go for. I expect a lot of demand for them!

There was a post recently, that showed the DGX spark is twice as fast as the M3 Ultra at prompt processing, but half as fast at output tokens [2]. They used gps-oss-120 for that test with a small context.

[1] https://gpuquicklist.com/apus?models=GB10%20Grace%20Blackwel...

[2] https://aimultiple.com/dgx-spark-alternatives , https://news.ycombinator.com/item?id=48732679

I’ve got a 128gb m5 max mbp and two sparks. For my real-world use cases, a single spark running DS4 Flash will have fully responded by the time my Mac has even started generating tokens. I figured I’d have more generation heavy work when I also bought the Mac, but it has done very little work running LLMs since I got the first Spark.

I mostly run them clustered for DS4 and am quite happy with the performance, and the cost isn’t that much more for two than the MBP while giving me double the unified memory.

I’ll probably pick up a third to run multiple smaller models. I don’t understand why people would buy a halo over a spark at comparable prices, particularly because if you want to cluster, the cx7 be beats the shit out of them when it comes to latency and throughput

Apple already announced that will increase prices.. doubt that Mac studio will stay cheap
MacBook Pro can come configured with 128GB M5 Max.
Apple rumor mill is suggesting that we may see M5 Mac Studio announced in September at that event. Apple just increased their pricing across the board, so I'm not holding my breath that they will be reasonably priced. They also cancelled development of their M6 in favor of the future M7 which I suspect will be a major, AI-focused upgrade compared to the M4->M5 upgrade
They cancelled the higher-end M6 models (including the Pro, this time).

There will be a basic M6.

Yeah, folks should be aware that if you're filling up the memory on a Strix Halo for an inference workload, you're going to be getting uncomfortably slow token rates. Like, DS4 (a 1-bit quantization of DeepSeek V4 Flash) runs at something like 9-13 tokens/second, with a loooong time to first token. It is not a realistic interactive coding model for agentic use.

I like my Strix Halo and keep it chewing on stuff, mostly non-interactive workloads (security audits of software mostly, training experiments, etc.), I get a lot of use out of it. If you want to experiment with AI, it is a good platform for that, though at $4k you can get an Nvidia-based Asus Ascend GX10, which is probably better. But, if you want a local model for interactive agentic use, you're going to be running either Qwen 3.6 or Gemma 4, which will fit comfortably on 2x64GB GPUs (even old GPUs will run them faster than the Strix Halo...I have dual Radeon Pro V620s which are faster, and they're six years old), or snugly on 32GB. A 48GB or 64GB Mac would run them well. Two Radeon AI Pro R9700 GPUs is probably the sweet spot, right now for GPUs. Not the cost of a good used car, like a 5090 or 4090, but plenty of memory and performance for local inference. Also, not finicky and weird and needing custom 3D printed fan shrouds like the old server GPUs on eBay.

At the moment, there just isn't a model that works better on a 128GB inference machine like this that don't also work fine on 64GB machines, which may be faster (very few 32GB GPUs will be slower, though I wouldn't recommend buying any GPU that isn't currently actively supported by the vendor drivers and CUDA or ROCm...so probably don't buy an MI50 or V100 or whatever).

this is not true. q2 Deepseek flash works on AMD Strix halo with pretty good results. Benchmark except:

ctx_tokens,prefill_tokens,prefill_tps,gen_tokens,gen_tps,kvcache_bytes 2048,2048,202.02,128,15.31,52184460 4096,2048,211.03,128,14.64,80373132 6144,2048,208.04,128,14.59,108561804 8192,2048,200.78,128,14.43,136750476 10240,2048,203.04,128,14.37,164939148 12288,2048,200.82,128,14.27,193127820 14336,2048,198.62,128,14.22,221316492 16384,2048,196.14,128,14.20,249505164 18432,2048,189.48,128,14.13,277693836 20480,2048,186.59,128,14.06,305882508 22528,2048,183.88,128,13.99,334071180 24576,2048,183.38,128,13.92,362259852 26624,2048,181.57,128,13.87,390448524 28672,2048,183.46,128,13.80,418637196 30720,2048,181.80,128,13.73,446825868 32768,2048,175.93,128,13.55,475014540 34816,2048,175.42,128,13.46,503203212

https://kyuz0.github.io/strix-halo-ds4-toolbox/

I said, "Like, DS4 (a 1-bit quantization of DeepSeek V4 Flash) runs at something like 9-13 tokens/second, with a loooong time to first token."

Which almost exactly matches the benchmark you just linked. Looks like it's possible to goose it to 15 tokens per second with a tiny context, but why would I want a giant model with a 2k context? DeepSeek is too big to be fast enough on a Strix Halo.

To be clear, if you think that's comfortable for interactive use, more power to you. But, I'm not waiting for that. I'll pay DeepSeek to host it for me. Their token prices are quite cheap and their cached tokens are even cheaper...and they have the most effective caching in the business, as far as I can tell. Even naively using the API you get 80-90% cached token rate. If you use Reasonix, you get ~98% cached token rate. I just built a feature for an app I'm working on for $0.10 for 20 minutes of work. Not bad at all.

At the moment, for around $4000-5000 you can either have speed (a GPU + 32GB VRAM), or you can have capacity - a DGX Spark/Halo, but not both.

I think once someone comes up with a machine which has both it will easily sell for $10000 and people will be queueing to buy it.

That’s about the price of a workstation-class Nvidia GPU nowadays and if you splurge double, you can get a PCIe version of a data center card.
Sure but where do you stop spending? :)

https://github.com/jamesob/local-llm

Until your agents have their own PCs and a little army of agents. :)
To be clear though that's GB/s. Which is 2 terabits/sec
And 4000 USD is over 1.2 million Hungarian forints.
yea but 4000 USD dont get confused for Hungarian forints just because one says 4000 usd instead.