|
|
|
|
|
by Alpha3031
11 days ago
|
|
Apple does have its own models for on-device use and the newest series has a MoE with expert activation swapped per prompt instead of per token which is interesting. I'd assume it's probably somewhat based off Gemini distillation or something, and I'm not expecting them to release the weights this time at least, but it would definitely be interesting to see an open local model using a similar MoE whether that's at a similar size (20B A1–4B supposedly) or a larger model (something DeepSeek V4 Flash sized, or slightly smaller, would just barely fit on consumer hardware, and could be fun to see how it performs I think). |
|