Y
Hacker News
new
|
ask
|
show
|
jobs
by
sergiotapia
7 days ago
Cerebras does not share the quantization of the models so you don't know if you're getting real K3 or k3 lite or something else.
1 comments
codexon
7 days ago
It most likely will be quantized. A cerebras wafer only has 44gb ram, and linking them together vastly reduces the speedup.
link