|
|
|
|
|
by hirako2000
149 days ago
|
|
Also because it's a large PR. Also because the maintainer has better things to do than taking longer and more energy to review than the author spent to write it, just to find that multiple optimisations will be requested, which the author may not be able to take on. the creator of llama.cc can hardly be suspected to be reluctant or biased towards GenAI. |
|
I wanted to see if Claude Code could port the HF / MLX implementation to llama.cpp and it was successful -- in my mind that's wild!
I also learned a ton about GPU programming, how omni models work, and refined my approach to planning large projects with automated end to end integration tests.
The PR was mostly to let people know about the code and weights, since there are quite a few comments requesting support:
https://github.com/ggml-org/llama.cpp/issues/16186