Hacker News new | ask | show | jobs
by Melatonic 30 days ago
Sounds like the perfect use case for some kind of framework where you have a local LLM (that can run on lower spec hardware) collaborating with the main LLM to optimise latency and all the other niche and legacy use cases ?